Senior Site Reliability Engineer – AI Infrastructure

Stelle nicht auf LinkedIn

🕒 vor 4 Monaten

🏄 California – Remote

infoinfo

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

👻 Geisterscore 37%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Andromeda

Andromeda

11 - 50 Mitarbeiter

🏥 Gesundheitswesen

💼 Beratung

🏨 Gastgewerbe

🔥 Finanzierung im letzten Jahr

💰 €15.142.238 Series A - Andromeda Robotics im 2025-09

Healthcare • Consulting • Hospitality

Andromeda ist ein GPU-Computing-Service und Marktplatz, der sofortigen Zugriff auf große Cluster von H100-, H200- und B200-Beschleunigern für Experimente, umfassendes Training und Inferenz bietet. Er unterstützt die Orchestrierung mit Slurm, Kubernetes oder direktem SSH, bietet flexible Nutzung ohne Mindestdauer zu wettbewerbsfähigen Preisen und umfasst DevOps-Expertise sowie lokales NAS- oder gestreamtes Speichern ohne Eingangs-/Ausgangsgebühren und 24/7-Support mit Branchen-SLAs. Das Unternehmen betreibt außerdem einen Drittanbietermarkt für GPUs unter gpulist. ai.

Beschreibung

• Design and evolve multi-provider, multi-region GPU compute clusters optimized for large-scale training • Serve as the primary technical point of contact for customers running large-scale training workloads • Define SLOs and error budgets that account for the unique failure modes of GPU infrastructure • Ensure the health and performance of high-speed interconnects • Build deep visibility into GPU utilization, memory pressure, interconnect throughput • Build production-grade automation for cluster provisioning, GPU health checks, job scheduling • Lead incident response for complex failures spanning hardware, networking, orchestration

🎯 Anforderungen

• Deep, hands-on experience operating large-scale GPU clusters (NVIDIA A100/H100/B200 or equivalent) • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training • Working knowledge of NCCL, CUDA, PyTorch distributed, DeepSpeed, Megatron, FSDP, or similar • Expert-level Linux knowledge • Strong experience running Kubernetes in production with GPU workloads • Strong engineering skills in Python, Go, or Bash • Hands-on experience building monitoring and alerting for GPU infrastructure • Proven track record leading incident response for complex distributed systems

🏖️ Vorteile

• Health insurance • Retirement plans • Paid time off • Flexible work arrangements • Professional development

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 4 Monaten

EITACIES Inc.

51 - 200

💼 Beratung

🏥 Gesundheitswesen

🏭 Fertigung

DevOps Architect leading platform engineering standards across a multi-cloud, hybrid environment at Eitacies Inc. Focus on automation, infrastructure, and cloud architecture.

🇺🇸 Vereinigte Staaten – Remote

💵 $60 / Stunde

⏰ Vollzeit

🟠 Senior

🔴 Experte

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Monaten

Avive Solutions Inc.

11 - 50

🏥 Gesundheitswesen

💼 Beratung

📦 Logistik

DevOps Engineer for Avive Solutions, building cloud infrastructure to revolutionize cardiac arrest responses. Collaborate cross-functionally to optimize systems for high-impact healthcare technology.

🇺🇸 Vereinigte Staaten – Remote

💵 $140.000 - $180.000 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Monaten

Runlayer

11 - 50

🤖 Künstliche Intelligenz

🔒 Cybersecurity

☁️ SaaS

Site Reliability Engineer ensuring performance and scalability of Runlayer’s AI infrastructure. Collaborating with founders and engineers in a fast-paced environment to support cloud and on-prem setups.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Monaten

Codvo.ai

51 - 200

🤖 Künstliche Intelligenz

🔒 Cybersecurity

☁️ SaaS

DevOps Engineer overseeing 24/7 support operations in a global tech services company. Leading a team and implementing automation for seamless application performance.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Monaten

Mediastream

51 - 200

💼 Beratung

📣 Marketing

📦 Logistik

DevOps Engineer at GlobalBet creating virtual sports solutions and managing Kubernetes clusters. Collaborating with a dynamic team to develop new game experiences.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich