Site Reliability Engineer – AI & ML Infrastructure, Kubernetes, AWS, Terraform

Vaga não está no LinkedIn

🕒 Março 10

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $220.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

👻 Score fantasma 41%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Deepgram

Deepgram

51 - 200 funcionários

Fundada em 2015

💼 Consultoria

🏥 Saúde

📦 Logística

💰 $47.000.000 Series B em 2022-11

Consulting • Healthcare • Logistics

A Deepgram é uma empresa líder em IA de voz que fornece APIs poderosas para aplicações de reconhecimento de fala, síntese de texto para fala e entendimento de linguagem. Sua plataforma permite que os desenvolvedores criem soluções avançadas de IA de voz para casos de uso como centrais de atendimento, transcrição médica, IA conversacional, entre outros. Conhecida por sua precisão inigualável, velocidade e custo-benefício, a tecnologia da Deepgram é confiada por grandes empresas e startups em todo o mundo. Oferecendo capacidades de transcrição em tempo real e altamente precisas, a Deepgram ajuda as empresas a obter insights a partir de dados de voz, tornando-se uma ferramenta essencial para transformar interações de voz.

Descrição

• Architect and maintain our core computing platform using Kubernetes on AWS and on-premise, providing a stable, scalable environment for all applications and services. • Develop and manage our entire infrastructure using Infrastructure-as-Code (IaC) principles with Terraform, ensuring our environments are reproducible, versioned, and automated. • Design, build, and optimize our AI/ML job scheduling and orchestration systems, integrating Slurm with our Kubernetes clusters to efficiently manage GPU resources. • Provision, manage, and maintain our on-premise bare metal server infrastructure for high-performance GPU computing. • Implement and manage the platform's networking (CNI, service mesh) and storage (CSI, S3) solutions to support high-throughput, low-latency workloads across hybrid environments. • Develop a comprehensive observability stack (monitoring, logging, tracing) to ensure platform health, and create automation for operational tasks, incident response, and performance tuning. • Collaborate with AI researchers and ML engineers to understand their infrastructure needs and build the tools and workflows that accelerate their development cycle. • Automate the life cycle of single-tenant, managed deployments

🎯 Requisitos

• 5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering (SRE) • Proven, hands-on experience building and managing production infrastructure with Terraform • Expert-level knowledge of Kubernetes architecture and operations in a large-scale environment • Experience with high-performance compute (HPC) job schedulers, specifically Slurm, for managing GPU-intensive AI workloads • Experience managing bare metal infrastructure, including server provisioning (e.g., PXE boot, MAAS), configuration, and lifecycle management • Strong scripting and automation skills (e.g., Python, Go, Bash)

🏖️ Benefícios

• Medical, dental, vision benefits • Annual wellness stipend • Mental health support • Life, STD, LTD Income Insurance Plans • Unlimited PTO • Generous paid parental leave • Flexible schedule • 12 Paid US company holidays • Quarterly personal productivity stipend • One-time stipend for home office upgrades • 401(k) plan with company match • Tax Savings Programs • Learning / Education stipend • Participation in talks and conferences • Employee Resource Groups • AI enablement workshops / sessions

Candidatar-se

Vagas Similares

🕒 Março 7

Inetum

10.000+ funcionários

💼 Consultoria

🏥 Saúde

🛡️ Seguros

Expert DevOps / DevSecOps supporting Generative AI initiatives at Inetum for digital transformation in the United States. Designing high-value GenAI use cases and integrating new tools and practices.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 Post-IPO Equity em 2007-03

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇫🇷 Francês obrigatório

🕒 Março 4

Akamai Technologies

5001 - 10000

🔒 Cibersegurança

🏢 Corporativo

📱 Mídia

Senior II DevOps Engineer developing and maintaining cloud infrastructures and web applications for top-tier security solutions. Engaging with highly skilled colleagues in a dynamic learning environment.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $112.500 - $202.500 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Fevereiro 27

Fuze Health

1001 - 5000

🏥 Saúde

☁️ SaaS

💊 Farmacêutico

Senior DevSecOps Engineer securing AWS/GCP infrastructure, Kubernetes, and CI/CD for Fuze Health’s nationwide digital pharmacy platform. Leading compliance, IAM, and incident resilience.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $128.000 - $160.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Fevereiro 27

Andromeda

11 - 50

🏥 Saúde

💼 Consultoria

🏨 Hospitalidade

Site Reliability Engineer managing Kubernetes-based clusters for AI infrastructure company Andromeda. Building reliable and scalable AI systems while working directly with customers and engineering teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $15.142.238 Series A - Andromeda Robotics em 2025-09

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Fevereiro 25

WorkOS

51 - 200

🔌 API

🏢 Corporativo

🤝 B2B

Site Reliability Engineer ensuring reliability and performance at WorkOS across complex systems. Leading incident response and collaborating with cross-functional teams for operational excellence.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $175.000 - $275.000 / ano

💰 $80.000.000 Series B - WorkOS em 2022-05

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório