Site Reliability Engineer – AI Infrastructure

Vaga não está no LinkedIn

🕒 Fevereiro 27

🏄 California – Remoto

info

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Andromeda

Andromeda

11 - 50 funcionários

🏥 Saúde

💼 Consultoria

🏨 Hospitalidade

🔥 Investimento no último ano

💰 $15.142.238 Series A - Andromeda Robotics em 2025-09

Healthcare • Consulting • Hospitality

A Andromeda é um serviço de computação com GPU e um marketplace que oferece acesso instantâneo a grandes clusters de aceleradores H100, H200 e B200 para experimentos, treinamentos em larga escala e inferência. Suporta orquestração com Slurm, Kubernetes ou SSH direto, oferece uso flexível sem duração mínima e preços competitivos, inclui expertise em DevOps, armazenamento local NAS ou transmitido sem taxas de entrada/saída, e suporte 24/7 com SLAs do setor. A empresa também opera um marketplace de GPUs de terceiros em gpulist. ai.

Descrição

• Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers • Build automation and tooling to streamline cluster deployments and integrations • Debug customer issues across networking, storage, scheduling, and system layers • Improve reliability and scalability of both training and inference infrastructure • Design and implement monitoring, alerting, and observability for critical systems • Collaborate with engineering and product teams to plan and deliver infrastructure for new services • Participate in on-call and incident response, leading postmortems and reliability improvements

🎯 Requisitos

• 5+ years experience in SRE, DevOps, or infrastructure engineering roles • Strong Linux systems and networking fundamentals • Deep experience with Kubernetes and container orchestration at scale • Proficiency with Infrastructure-as-Code (Terraform, Helm, Ansible, etc.) • Strong automation and scripting skills (Python, Go, or Bash) • Experience with observability stacks (Prometheus, Grafana, Loki, Datadog, etc.) • Track record of operating production systems and leading incident response

🏖️ Benefícios

• Ownership and autonomy to shape systems • Opportunities to work directly with customers and providers

Candidatar-se

Vagas Similares

🕒 Fevereiro 25

Nick AI

1 - 10

💼 Consultoria

📦 Logística

🤖 Inteligência Artificial

Backend/DevOps Engineer managing deployments and infrastructure for AI trading platform. Responsible for security, reliability, and scaling of systems across multiple venues.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Fevereiro 25

WorkOS

51 - 200

🔌 API

🏢 Corporativo

🤝 B2B

Site Reliability Engineer ensuring reliability and performance at WorkOS across complex systems. Leading incident response and collaborating with cross-functional teams for operational excellence.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $175.000 - $275.000 / ano

💰 $80.000.000 Series B - WorkOS em 2022-05

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Fevereiro 25

Vultr

201 - 500

🤖 Inteligência Artificial

🤝 B2B

🔧 Hardware

NetDevOps Engineer for RDMA Fabric Automation at Vultr. Automating and operating Ethernet fabrics with a focus on network performance.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $90.000 - $130.000 / ano

💰 $329.000.000 Debt Financing - Vultr em 2025-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Fevereiro 25

Tactibit Technologies

11 - 50

💼 Consultoria

📦 Logística

🎖️ Defesa

DevOps Engineer working at Tactibit Technologies to modernize legacy architectures for mission-critical systems. Collaborate with teams on cloud migrations and automating business processes.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Fevereiro 25

DroneUp

51 - 200

💼 Consultoria

📦 Logística

🏭 Manufatura

SRE - Platform Engineer at DroneUp focusing on IT infrastructure reliability and scalability. Driving SRE best practices within the team and collaborating on cloud engineering solutions.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $125.000 - $150.000 / ano

💰 $241.201 Seed Round - DroneUp em 2022-07

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório