Senior Site Reliability Engineer – AI Infrastructure

Vaga não está no LinkedIn

🕒 Abril 9

🏄 California – Remoto

infoinfo

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

👻 Score fantasma 43%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Andromeda

Andromeda

11 - 50 funcionários

🏥 Saúde

💼 Consultoria

🏨 Hospitalidade

💰 $15.142.238 Series A - Andromeda Robotics em 2025-09

Healthcare • Consulting • Hospitality

A Andromeda é um serviço de computação com GPU e um marketplace que oferece acesso instantâneo a grandes clusters de aceleradores H100, H200 e B200 para experimentos, treinamentos em larga escala e inferência. Suporta orquestração com Slurm, Kubernetes ou SSH direto, oferece uso flexível sem duração mínima e preços competitivos, inclui expertise em DevOps, armazenamento local NAS ou transmitido sem taxas de entrada/saída, e suporte 24/7 com SLAs do setor. A empresa também opera um marketplace de GPUs de terceiros em gpulist. ai.

Descrição

• Design and evolve multi-provider, multi-region GPU compute clusters optimized for large-scale training • Serve as the primary technical point of contact for customers running large-scale training workloads • Define SLOs and error budgets that account for the unique failure modes of GPU infrastructure • Ensure the health and performance of high-speed interconnects • Build deep visibility into GPU utilization, memory pressure, interconnect throughput • Build production-grade automation for cluster provisioning, GPU health checks, job scheduling • Lead incident response for complex failures spanning hardware, networking, orchestration

🎯 Requisitos

• Deep, hands-on experience operating large-scale GPU clusters (NVIDIA A100/H100/B200 or equivalent) • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training • Working knowledge of NCCL, CUDA, PyTorch distributed, DeepSpeed, Megatron, FSDP, or similar • Expert-level Linux knowledge • Strong experience running Kubernetes in production with GPU workloads • Strong engineering skills in Python, Go, or Bash • Hands-on experience building monitoring and alerting for GPU infrastructure • Proven track record leading incident response for complex distributed systems

🏖️ Benefícios

• Health insurance • Retirement plans • Paid time off • Flexible work arrangements • Professional development

Candidatar-se

Vagas Similares

🕒 Abril 8

EITACIES Inc.

51 - 200

💼 Consultoria

🏥 Saúde

🏭 Manufatura

DevOps Architect leading platform engineering standards across a multi-cloud, hybrid environment at Eitacies Inc. Focus on automation, infrastructure, and cloud architecture.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $60 / hora

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Abril 3

Runlayer

11 - 50

🤖 Inteligência Artificial

🔒 Cibersegurança

☁️ SaaS

Site Reliability Engineer ensuring performance and scalability of Runlayer’s AI infrastructure. Collaborating with founders and engineers in a fast-paced environment to support cloud and on-prem setups.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Abril 2

Mediastream

51 - 200

💼 Consultoria

📣 Marketing

📦 Logística

DevOps Engineer at GlobalBet creating virtual sports solutions and managing Kubernetes clusters. Collaborating with a dynamic team to develop new game experiences.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Abril 1

NVision IT, LLC

11 - 50

🎖️ Defesa

🏥 Saúde

📦 Logística

DevOps consultant at NVision IT supporting federal projects with technical and strategic server support. Requires U.S. citizenship and active security clearance.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Abril 1

MLabs

51 - 200

💼 Consultoria

🤖 Inteligência Artificial

💳 Fintech

Senior DevOps / SRE Engineer working on infrastructure for autonomous AI trading agents. Managing high-stakes environments for real-time financial workloads with robust reliability.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $120.000 - $150.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório