Forward Deployed Engineer – SRE

Stelle nicht auf LinkedIn

🕒 vor 1 Monat

🏄 California – Remote

infoinfo

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

👻 Geisterscore 10%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Andromeda

Andromeda

11 - 50 Mitarbeiter

🏥 Gesundheitswesen

💼 Beratung

🏨 Gastgewerbe

🔥 Finanzierung im letzten Jahr

💰 €15.142.238 Series A - Andromeda Robotics im 2025-09

Healthcare • Consulting • Hospitality

Andromeda ist ein GPU-Computing-Service und Marktplatz, der sofortigen Zugriff auf große Cluster von H100-, H200- und B200-Beschleunigern für Experimente, umfassendes Training und Inferenz bietet. Er unterstützt die Orchestrierung mit Slurm, Kubernetes oder direktem SSH, bietet flexible Nutzung ohne Mindestdauer zu wettbewerbsfähigen Preisen und umfasst DevOps-Expertise sowie lokales NAS- oder gestreamtes Speichern ohne Eingangs-/Ausgangsgebühren und 24/7-Support mit Branchen-SLAs. Das Unternehmen betreibt außerdem einen Drittanbietermarkt für GPUs unter gpulist. ai.

Beschreibung

• Serve as the primary technical point of contact for teams running large-scale training and inference workloads • Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns • Profile and improve distributed training performance on live workloads • Own reliability outcomes for the accounts you're deployed on • Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) • Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks • Turn every repeated deployment problem into automation

🎯 Anforderungen

• Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent) • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training • Working knowledge of how large training and inference jobs actually run • Expert-level Linux experience: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at the syscall and hardware level • Strong experience running Kubernetes in production with GPU workloads • Strong engineering skills in Python, Go, or Bash • Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent) • Hands-on experience building monitoring and alerting for GPU-specific telemetry

🏖️ Vorteile

• Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage • 401(k) • Unlimited PTO

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Monat

Smithfield Foods

10.000+ Mitarbeiter

🏭 Fertigung

🌾 Landwirtschaft

🍽️ Lebensmittel & Getränke

Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

CI&T

5001 - 10000

💼 Beratung

🏥 Gesundheitswesen

📣 Marketing

Lead design, deployment, and optimization of NVIDIA cuOpt solutions for AI-focused business operations. Collaborate across teams for effective routing and optimization solutions in Colombia.

🇺🇸 Vereinigte Staaten – Remote

💰 €5.500.000 Venture Round im 2014-04

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Rocket.net

11 - 50

☁️ SaaS

🛍️ eCommerce

🏢 Unternehmen

Site Reliability Engineer ensuring high standards for servers, services, and customer environments at Rocket.net. Providing advanced technical support and ensuring platform reliability.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Neural Earth

11 - 50

🤖 Künstliche Intelligenz

🛡️ Versicherung

🏠 Immobilien

DevOps Lead at Neural Earth responsible for maintaining CI/CD pipelines and monitoring AWS cloud infrastructure. Support incident response and operational work for smooth engineering processes.

🇺🇸 Vereinigte Staaten – Remote

💵 $125.000 - $156.000 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

PerfectServe

201 - 500

🏥 Gesundheitswesen

⚕️ Krankenversicherung

☁️ SaaS

Forward Deployment Engineer, AI optimizing and deploying AI voice agent solutions for PerfectServe's customers. Collaborating with Product and Customer Success to ensure successful go-live and ongoing support.

🇺🇸 Vereinigte Staaten – Remote

💵 $120.000 - $140.000 / Jahr

💰 Private Equity Round im 2018-05

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich