Site Reliability Engineer

Stelle nicht auf LinkedIn

🕒 vor 1 Monat

🤠 Texas – Remote

infoinfo

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

👻 Geisterscore 16%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of STN Incorporated

STN Incorporated

11 - 50 Mitarbeiter

Gegründet 2016

🏢 Unternehmen

🔒 Cybersecurity

🔧 Hardware

Enterprise • Cybersecurity • Hardware

STN Incorporated ist ein Anbieter für IT- und Cloud-Infrastruktur auf Unternehmensniveau, der sichere, prüfbereite Infrastrukturen für geschäftskritische Systeme und anspruchsvolle KI-Arbeitslasten liefert. STN betreibt ein Managed-Operating-Modell, das private CPU-Clouds, GPU One AI-Infrastruktur, sichere Netzwerke und Speicher sowie 24/7 menschlichen Support mit hohen Verfügbarkeits-SLAs bietet. Zu den Dienstleistungen gehören verwaltete Infrastruktur- und Cloud-Operationen, Cybersecurity-Operationen und Vorfallreaktionen, Compliance- und Risikomanagement (SOC 2 Typ II, HIPAA-konform), Backup und Wiederherstellung sowie Beschaffung und Lebenszyklusmanagement von Unternehmens-Technologien. STN bedient Unternehmen, wachstumsstarke SaaS-Unternehmen, KI-Entwickler und Modellbauer, Robotik-/physische KI-Firmen und regulierte Branchen wie das Gesundheitswesen.

Beschreibung

• Define and operate Service Level Objectives (SLOs) aligned with customer SLAs • Build and maintain the observability stack including metrics, logs, traces, and alerting • Lead incident response and chair post-incident reviews • Drive automation to reduce toil and improve mean-time-to-recover (MTTR) • Author and maintain operational runbooks alongside the NOC • Manage on-call rotation, escalation paths, and incident-management tooling • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering • Drive chaos engineering, game days, and reliability testing programs • Produce SLA performance reports in coordination with the SLA Manager • Mentor junior engineers and contribute to engineering culture

🎯 Anforderungen

• 5+ years in SRE, DevOps, or production engineering roles • Strong programming skills in Go, Python, or both • Hands-on experience operating Kubernetes-based platforms at scale • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry) • Strong incident management experience including major-incident command

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Monat

Zafran Security

51 - 200

🔐 Sicherheit

Senior DevOps Engineer at Zafran working on compliance certifications and implementing security controls across infrastructure. Collaborating with teams to enhance security posture and ensure regulatory compliance.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Runpod

51 - 200

🤖 Künstliche Intelligenz

☁️ SaaS

🤝 B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

🇺🇸 Vereinigte Staaten – Remote

💵 $150.000 - $200.000 / Jahr

💰 €20.000.000 Seed im 2024-06

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Multi Media, LLC

51 - 200

💼 Beratung

📣 Marketing

📱 Medien

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

🇺🇸 Vereinigte Staaten – Remote

💵 $169.000 - $215.000 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Thumbtack

1001 - 5000

🏪 Marktplatz

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 Vereinigte Staaten – Remote

💵 $179.400 - $232.100 / Jahr

💰 €75.000.000 Debt Financing - Thumbtack im 2024-07

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

ICF

5001 - 10000

💼 Beratung

🏛️ Regierung

🏥 Gesundheitswesen

DevOps Engineer building healthcare reporting services for ICF. Implementing cloud-based solutions using AWS and fostering collaboration on CI/CD pipeline improvements.

🇺🇸 Vereinigte Staaten – Remote

💵 $108.476 - $184.409 / Jahr

💰 €29.000.000 Grant im 2023-03

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich