Senior Site Reliability Engineer

🕒 vor 4 Monaten

🍂 Massachusetts – Remote

infoinfo

💵 $121.400 - $218.600 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

👻 Geisterscore 41%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Akamai Technologies

Akamai Technologies

5001 - 10000 Mitarbeiter

🔒 Cybersecurity

🏢 Unternehmen

📱 Medien

Cybersecurity • Enterprise • Media

Akamai Technologies ist eine globale Plattform für Edge- und Cloud-Dienstleistungen, die Lösungen für die Bereitstellung von Inhalten, Edge Computing und Sicherheit anbietet. Das Unternehmen betreibt eines der weltweit größten verteilten Netzwerke zur Beschleunigung und zum Schutz von Web-, Medien- und Anwendungsverkehr. Akamai bietet Produkte für die Bereitstellung von Inhalten, DDoS-Schutz, API- und App-Sicherheit, Bot-Management, Edge Computing (serverlose/Edge-Funktionen) und KI-Inferenz am Edge an. Zudem stellt Akamai unternehmensfokussierte Sicherheitsdienste bereit (Zero Trust, Identitäts- und Zugangsmanagement, sicherer Internetzugang) sowie Tools für Cloud-/KI-Infrastruktur. Kürzlich hat Akamai seine Fähigkeiten durch Akquisitionen (zum Beispiel LayerX) erweitert, um KI-Nutzungen im Browser zu steuern.

Beschreibung

• Own reliability workstreams for Akamai's serverless inference platform • Build automation and tooling • Contribute to architecture and operational decisions • Take ownership of critical reliability problems end-to-end • Partner with product engineering teams • Develop expertise in GPU infrastructure, Kubernetes at scale, and AI inference workloads • Build and maintain observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking • Write automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response • Integrate AI workloads into Akamai's incident management processes • Build and maintain CI/CD integrations, deployment safety checks, and rollback automation • Collaborate with product engineering teams to improve reliability and ensure operational readiness for product releases • Contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure

🎯 Anforderungen

• 5+ years of experience in SRE, infrastructure engineering, or platform engineering, working with large-scale distributed systems • Extensive experience with Kubernetes and containerization at scale • Experience defining SLOs and working with observability tools such as Prometheus, Grafana, and distributed tracing • Coding ability in Python or Go for automation and tooling, with experience in CI/CD pipelines, deployment safety, and infrastructure-as-code • Interest in or experience with AI/ML infrastructure, model serving, or GPU workloads • Ability to take ownership of problems and drive them to resolution independently

🏖️ Vorteile

• Healthcare • 401K savings plan • Company holidays • Vacation (in the form of PTO) • Sick time • Family friendly benefits including parental leave • Employee assistance program focusing on mental and financial wellness • Flexible working arrangements

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 4 Monaten

Red River

501 - 1000

💼 Beratung

📦 Logistik

Senior Wireless Deployment Engineer managing deployment, optimization, and lifecycle of enterprise wireless networks. Providing technical leadership and support for Aruba and Juniper Mist solutions.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

SGNL

11 - 50

🔒 Cybersecurity

🔐 Sicherheit

☁️ SaaS

Senior DevOps Engineer at SGNL solving authorization challenges for major companies. Collaborating and leading teams in a dynamic, scale-oriented environment.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

Cority

201 - 500

🏥 Gesundheitswesen

📦 Logistik

💼 Beratung

Sr. DevOps Engineer working to deploy and operate systems at Cority, the global EHS software provider. Collaborating with engineering for continuous delivery and monitoring towards operational excellence.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

Mark43

201 - 500

🤖 Künstliche Intelligenz

🏛️ Regierung

🔒 Cybersecurity

DevOps Engineer focusing on improving platform reliability and developer workflows at Mark43. Leading initiatives to enhance CI/CD and operational metrics while collaborating across teams.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

Orbital Engineering, Inc.

501 - 1000

🏗️ Bauwesen

💼 Beratung

🏭 Fertigung

Electrical Reliability Engineer managing risk-based electrical infrastructure programs across the U.S. for Orbital, focusing on improving system reliability and minimizing outages.

🇺🇸 Vereinigte Staaten – Remote

💵 $125.000 - $175.000 / Jahr

⏰ Vollzeit

🟠 Senior

🔴 Experte

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich