
11 - 50 Mitarbeiter
🏥 Gesundheitswesen
💼 Beratung
🏨 Gastgewerbe
🔥 Finanzierung im letzten Jahr
💰 €15.142.238 Series A - Andromeda Robotics im 2025-09
Healthcare • Consulting • Hospitality
Andromeda ist ein GPU-Computing-Service und Marktplatz, der sofortigen Zugriff auf große Cluster von H100-, H200- und B200-Beschleunigern für Experimente, umfassendes Training und Inferenz bietet. Er unterstützt die Orchestrierung mit Slurm, Kubernetes oder direktem SSH, bietet flexible Nutzung ohne Mindestdauer zu wettbewerbsfähigen Preisen und umfasst DevOps-Expertise sowie lokales NAS- oder gestreamtes Speichern ohne Eingangs-/Ausgangsgebühren und 24/7-Support mit Branchen-SLAs. Das Unternehmen betreibt außerdem einen Drittanbietermarkt für GPUs unter gpulist. ai.
🕒 vor 1 Monat
🏄 California – Remote
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
🦅 H1B-Visum-Sponsor
👻 Geisterscore 10%
🗣️🇺🇸🇬🇧 Englisch erforderlich
Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

11 - 50 Mitarbeiter
🏥 Gesundheitswesen
💼 Beratung
🏨 Gastgewerbe
🔥 Finanzierung im letzten Jahr
💰 €15.142.238 Series A - Andromeda Robotics im 2025-09
Healthcare • Consulting • Hospitality
Andromeda ist ein GPU-Computing-Service und Marktplatz, der sofortigen Zugriff auf große Cluster von H100-, H200- und B200-Beschleunigern für Experimente, umfassendes Training und Inferenz bietet. Er unterstützt die Orchestrierung mit Slurm, Kubernetes oder direktem SSH, bietet flexible Nutzung ohne Mindestdauer zu wettbewerbsfähigen Preisen und umfasst DevOps-Expertise sowie lokales NAS- oder gestreamtes Speichern ohne Eingangs-/Ausgangsgebühren und 24/7-Support mit Branchen-SLAs. Das Unternehmen betreibt außerdem einen Drittanbietermarkt für GPUs unter gpulist. ai.
• Serve as the primary technical point of contact for teams running large-scale training and inference workloads • Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns • Profile and improve distributed training performance on live workloads • Own reliability outcomes for the accounts you're deployed on • Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) • Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks • Turn every repeated deployment problem into automation
• Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent) • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training • Working knowledge of how large training and inference jobs actually run • Expert-level Linux experience: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at the syscall and hardware level • Strong experience running Kubernetes in production with GPU workloads • Strong engineering skills in Python, Go, or Bash • Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent) • Hands-on experience building monitoring and alerting for GPU-specific telemetry
• Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage • 401(k) • Unlimited PTO
Jetzt Bewerben🕒 vor 1 Monat
Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.
🇺🇸 Vereinigte Staaten – Remote
💵 $85.000 - $120.000 / Jahr
⏰ Vollzeit
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
🦅 H1B-Visum-Sponsor
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 1 Monat
Lead design, deployment, and optimization of NVIDIA cuOpt solutions for AI-focused business operations. Collaborate across teams for effective routing and optimization solutions in Colombia.
🇺🇸 Vereinigte Staaten – Remote
💰 €5.500.000 Venture Round im 2014-04
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 1 Monat
Site Reliability Engineer ensuring high standards for servers, services, and customer environments at Rocket.net. Providing advanced technical support and ensuring platform reliability.
🇺🇸 Vereinigte Staaten – Remote
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 1 Monat
DevOps Lead at Neural Earth responsible for maintaining CI/CD pipelines and monitoring AWS cloud infrastructure. Support incident response and operational work for smooth engineering processes.
🇺🇸 Vereinigte Staaten – Remote
💵 $125.000 - $156.000 / Jahr
⏰ Vollzeit
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 1 Monat
Forward Deployment Engineer, AI optimizing and deploying AI voice agent solutions for PerfectServe's customers. Collaborating with Product and Customer Success to ensure successful go-live and ongoing support.
🇺🇸 Vereinigte Staaten – Remote
💵 $120.000 - $140.000 / Jahr
💰 Private Equity Round im 2018-05
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
🗣️🇺🇸🇬🇧 Englisch erforderlich