Site Reliability Engineer – AI Infrastructure

Stelle nicht auf LinkedIn

🕒 vor 5 Monaten

🏄 California – Remote

infoinfo

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

👻 Geisterscore 44%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Andromeda

Andromeda

11 - 50 Mitarbeiter

🏥 Gesundheitswesen

💼 Beratung

🏨 Gastgewerbe

🔥 Finanzierung im letzten Jahr

💰 €15.142.238 Series A - Andromeda Robotics im 2025-09

Healthcare • Consulting • Hospitality

Andromeda ist ein GPU-Computing-Service und Marktplatz, der sofortigen Zugriff auf große Cluster von H100-, H200- und B200-Beschleunigern für Experimente, umfassendes Training und Inferenz bietet. Er unterstützt die Orchestrierung mit Slurm, Kubernetes oder direktem SSH, bietet flexible Nutzung ohne Mindestdauer zu wettbewerbsfähigen Preisen und umfasst DevOps-Expertise sowie lokales NAS- oder gestreamtes Speichern ohne Eingangs-/Ausgangsgebühren und 24/7-Support mit Branchen-SLAs. Das Unternehmen betreibt außerdem einen Drittanbietermarkt für GPUs unter gpulist. ai.

Beschreibung

• Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers • Build automation and tooling to streamline cluster deployments and integrations • Debug customer issues across networking, storage, scheduling, and system layers • Improve reliability and scalability of both training and inference infrastructure • Design and implement monitoring, alerting, and observability for critical systems • Collaborate with engineering and product teams to plan and deliver infrastructure for new services • Participate in on-call and incident response, leading postmortems and reliability improvements

🎯 Anforderungen

• 5+ years experience in SRE, DevOps, or infrastructure engineering roles • Strong Linux systems and networking fundamentals • Deep experience with Kubernetes and container orchestration at scale • Proficiency with Infrastructure-as-Code (Terraform, Helm, Ansible, etc.) • Strong automation and scripting skills (Python, Go, or Bash) • Experience with observability stacks (Prometheus, Grafana, Loki, Datadog, etc.) • Track record of operating production systems and leading incident response

🏖️ Vorteile

• Ownership and autonomy to shape systems • Opportunities to work directly with customers and providers

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 5 Monaten

WorkOS

51 - 200

🔌 API

🏢 Unternehmen

🤝 B2B

Site Reliability Engineer ensuring reliability and performance at WorkOS across complex systems. Leading incident response and collaborating with cross-functional teams for operational excellence.

🇺🇸 Vereinigte Staaten – Remote

💵 $175.000 - $275.000 / Jahr

💰 €80.000.000 Series B - WorkOS im 2022-05

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Monaten

Tactibit Technologies

11 - 50

💼 Beratung

📦 Logistik

🎖️ Verteidigung

DevOps Engineer working at Tactibit Technologies to modernize legacy architectures for mission-critical systems. Collaborate with teams on cloud migrations and automating business processes.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 6 Monaten

Filevine

201 - 500

☁️ SaaS

⚖️ Rechtswesen

🤖 Künstliche Intelligenz

Senior DBRE managing performance and scalability of data platform at Filevine, a legal AI company. Focus on AI-driven automation, optimizing SQL Server and Postgres environments.

🇺🇸 Vereinigte Staaten – Remote

💵 $145.000 - $180.000 / Jahr

💰 €108.000.000 Series D im 2022-04

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 6 Monaten

Ensono

1001 - 5000

💼 Beratung

DevOps Engineer working with AWS technologies on client deployment projects. Responsible for automation, support, and high availability of mission-critical solutions.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 6 Monaten

Ensono

1001 - 5000

💼 Beratung

Senior DevOps Engineer focused on deploying and supporting AWS technologies. Leading engineering projects and providing client support in a remote U.S. environment.

🗣️🇺🇸🇬🇧 Englisch erforderlich