Site Reliability Engineer

🕒 vor 1 Monat

🏄 California – Remote

infoinfo

💵 $150.000 - $200.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

👻 Geisterscore 16%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Runpod

Runpod

51 - 200 Mitarbeiter

Gegründet 2022

🤖 Künstliche Intelligenz

☁️ SaaS

🤝 B2B

💰 €20.000.000 Seed im 2024-06

Artificial Intelligence • SaaS • B2B

Runpod ist eine Cloud-Plattform, die bedarfsgerechte GPU-Rechenleistung und verwaltete Infrastruktur bietet, die speziell für die Entwicklung und Bereitstellung von KI optimiert sind. Sie bietet GPU-"Pods" in 31 globalen Regionen, serverlose GPU-Endpunkte für latenzarme Inferenz, Multi-Node-GPU-Cluster für verteiltes Training und ein Hub zur Bereitstellung von Open-Source-Modellen und -Vorlagen. Runpod legt Wert auf schnellen Start (unter 200 ms Cold Starts), automatisches Skalieren von null auf tausende Arbeiter, Unterstützung für über 30 GPU-SKUs und Werkzeuge für den kompletten KI-Lebenszyklus von Experimenten bis hin zur Produktion, wobei sie Entwickler und KI-Teams in Unternehmen fokussiert.

Beschreibung

• Define and implement SLIs/SLOs for critical services • Lead incident response and coordinate cross-team mitigation efforts • Conduct blameless postmortems and ensure corrective actions are completed • Perform production readiness reviews for new services and features • Identify systemic risks and drive preventative improvements • Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.) • Build internal tooling for reliability tracking and reporting • Automate recurring operational workflows • Strengthen CI/CD reliability and release processes • Partner with engineering teams to improve system resilience • Provide guidance on fault tolerance, scalability, and failure handling. • Contribute to architectural discussions with a reliability-first mindset.

🎯 Anforderungen

• 5+ years of experience in SRE, Reliability Engineering, or Production Engineering • Strong Linux systems and Networking expertise • Experience managing containerized production systems • Strong understanding of distributed systems and failure modes • Experience defining and managing SLIs/SLOs • Proven incident response and postmortem leadership experience • Strong scripting or programming skills • Experience with monitoring and alerting systems • Excellent written communication skills • Successful completion of a background check. • Preferred: Experience with GPU infrastructure or AI/ML platforms • Experience improving reliability in high-growth or large scale environments • Familiarity with GPU observability tooling • Experience with Infrastructure as Code • Experience working in startup environments • Experience building internal reliability platforms or frameworks.

🏖️ Vorteile

• Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. • Generous medical, dental & vision plans • Flexible PTO- take the time you need to recharge • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Monat

Multi Media, LLC

51 - 200

💼 Beratung

📣 Marketing

📱 Medien

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

🇺🇸 Vereinigte Staaten – Remote

💵 $169.000 - $215.000 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Thumbtack

1001 - 5000

🏪 Marktplatz

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 Vereinigte Staaten – Remote

💵 $179.400 - $232.100 / Jahr

💰 €75.000.000 Debt Financing - Thumbtack im 2024-07

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

ICF

5001 - 10000

💼 Beratung

🏛️ Regierung

🏥 Gesundheitswesen

DevOps Engineer building healthcare reporting services for ICF. Implementing cloud-based solutions using AWS and fostering collaboration on CI/CD pipeline improvements.

🇺🇸 Vereinigte Staaten – Remote

💵 $108.476 - $184.409 / Jahr

💰 €29.000.000 Grant im 2023-03

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

ICF

5001 - 10000

💼 Beratung

🏛️ Regierung

🏥 Gesundheitswesen

Senior DevOps Engineer delivering best in class healthcare reporting services for ICF. Working collaboratively to implement cloud solutions and establish CI/CD pipelines using AWS.

🇺🇸 Vereinigte Staaten – Remote

💵 $108.476 - $184.409 / Jahr

💰 €29.000.000 Grant im 2023-03

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Valence

51 - 200

🤖 Künstliche Intelligenz

👥 HR Tech

☁️ SaaS

Senior DevOps Engineer managing AWS infrastructure for a pioneering AI coaching platform. Leading security initiatives and collaborating with multiple development teams on scalable solutions.

🇺🇸 Vereinigte Staaten – Remote

🔥 Finanzierung im letzten Jahr

💰 €50.000.000 Series B - Valence im 2025-09

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich