Site Reliability Engineer – AI & ML Infrastructure, Kubernetes, AWS, Terraform

Stelle nicht auf LinkedIn

🕒 vor 5 Monaten

🇺🇸 Vereinigte Staaten – Remote

💵 $150.000 - $220.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

👻 Geisterscore 35%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Deepgram

Deepgram

51 - 200 Mitarbeiter

Gegründet 2015

💼 Beratung

🏥 Gesundheitswesen

📦 Logistik

💰 €47.000.000 Series B im 2022-11

Consulting • Healthcare • Logistics

Deepgram ist ein führendes Voice-AI-Unternehmen, das leistungsstarke APIs für Speech-to-Text, Text-to-Speech und Language-Understanding-Anwendungen bereitstellt. Die Plattform ermöglicht es Entwicklerinnen und Entwicklern, anspruchsvolle Voice-AI-Lösungen für Use Cases wie Contact Center, medizinische Transkription, Conversational AI und mehr zu entwickeln. Bekannt für unübertroffene Genauigkeit, Geschwindigkeit und Kosteneffizienz genießt die Technologie von Deepgram das Vertrauen führender Unternehmen und Start-ups weltweit. Mit Echtzeit- und hochgenauer Transkription hilft Deepgram Unternehmen, Erkenntnisse aus Sprachdaten zu gewinnen – ein unverzichtbares Werkzeug, um Sprachinteraktionen zu transformieren.

Beschreibung

• Architect and maintain our core computing platform using Kubernetes on AWS and on-premise, providing a stable, scalable environment for all applications and services. • Develop and manage our entire infrastructure using Infrastructure-as-Code (IaC) principles with Terraform, ensuring our environments are reproducible, versioned, and automated. • Design, build, and optimize our AI/ML job scheduling and orchestration systems, integrating Slurm with our Kubernetes clusters to efficiently manage GPU resources. • Provision, manage, and maintain our on-premise bare metal server infrastructure for high-performance GPU computing. • Implement and manage the platform's networking (CNI, service mesh) and storage (CSI, S3) solutions to support high-throughput, low-latency workloads across hybrid environments. • Develop a comprehensive observability stack (monitoring, logging, tracing) to ensure platform health, and create automation for operational tasks, incident response, and performance tuning. • Collaborate with AI researchers and ML engineers to understand their infrastructure needs and build the tools and workflows that accelerate their development cycle. • Automate the life cycle of single-tenant, managed deployments

🎯 Anforderungen

• 5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering (SRE) • Proven, hands-on experience building and managing production infrastructure with Terraform • Expert-level knowledge of Kubernetes architecture and operations in a large-scale environment • Experience with high-performance compute (HPC) job schedulers, specifically Slurm, for managing GPU-intensive AI workloads • Experience managing bare metal infrastructure, including server provisioning (e.g., PXE boot, MAAS), configuration, and lifecycle management • Strong scripting and automation skills (e.g., Python, Go, Bash)

🏖️ Vorteile

• Medical, dental, vision benefits • Annual wellness stipend • Mental health support • Life, STD, LTD Income Insurance Plans • Unlimited PTO • Generous paid parental leave • Flexible schedule • 12 Paid US company holidays • Quarterly personal productivity stipend • One-time stipend for home office upgrades • 401(k) plan with company match • Tax Savings Programs • Learning / Education stipend • Participation in talks and conferences • Employee Resource Groups • AI enablement workshops / sessions

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 5 Monaten

Inetum

10.000+ Mitarbeiter

💼 Beratung

🏥 Gesundheitswesen

🛡️ Versicherung

Expert DevOps / DevSecOps supporting Generative AI initiatives at Inetum for digital transformation in the United States. Designing high-value GenAI use cases and integrating new tools and practices.

🇺🇸 Vereinigte Staaten – Remote

💰 Post-IPO Equity im 2007-03

⏰ Vollzeit

🟠 Senior

🔴 Experte

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇫🇷 Französisch erforderlich

🕒 vor 5 Monaten

ALTEN Technology USA

501 - 1000

💼 Beratung

🎖️ Verteidigung

🏥 Gesundheitswesen

Design and Release Engineer developing vehicle components and systems from concept to production at ALTEN Technology USA.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Monaten

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

🏢 Unternehmen

📱 Medien

Senior II DevOps Engineer developing and maintaining cloud infrastructures and web applications for top-tier security solutions. Engaging with highly skilled colleagues in a dynamic learning environment.

🇺🇸 Vereinigte Staaten – Remote

💵 $112.500 - $202.500 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Monaten

Tiger Resourcing Group

11 - 50

🎖️ Verteidigung

💼 Beratung

📦 Logistik

DevOps Engineer for IT services seeking to design, build, and maintain infrastructure environments. Ensuring high-quality software delivery across internal and external networks with collaborative teamwork.

🇺🇸 Vereinigte Staaten – Remote

💵 $145.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Monaten

Andromeda

11 - 50

🏥 Gesundheitswesen

💼 Beratung

🏨 Gastgewerbe

Site Reliability Engineer managing Kubernetes-based clusters for AI infrastructure company Andromeda. Building reliable and scalable AI systems while working directly with customers and engineering teams.

🇺🇸 Vereinigte Staaten – Remote

🔥 Finanzierung im letzten Jahr

💰 €15.142.238 Series A - Andromeda Robotics im 2025-09

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich