Site Reliability Engineer – AI Infrastructure

Vaga não está no LinkedIn

🕒 Fevereiro 27

🏄 California – Remoto

infoinfo

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

👻 Score fantasma 45%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Andromeda

Andromeda

11 - 50 funcionários

🏥 Saúde

💼 Consultoria

🏨 Hospitalidade

💰 $15.142.238 Series A - Andromeda Robotics em 2025-09

Healthcare • Consulting • Hospitality

A Andromeda é um serviço de computação com GPU e um marketplace que oferece acesso instantâneo a grandes clusters de aceleradores H100, H200 e B200 para experimentos, treinamentos em larga escala e inferência. Suporta orquestração com Slurm, Kubernetes ou SSH direto, oferece uso flexível sem duração mínima e preços competitivos, inclui expertise em DevOps, armazenamento local NAS ou transmitido sem taxas de entrada/saída, e suporte 24/7 com SLAs do setor. A empresa também opera um marketplace de GPUs de terceiros em gpulist. ai.

Descrição

• Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers • Build automation and tooling to streamline cluster deployments and integrations • Debug customer issues across networking, storage, scheduling, and system layers • Improve reliability and scalability of both training and inference infrastructure • Design and implement monitoring, alerting, and observability for critical systems • Collaborate with engineering and product teams to plan and deliver infrastructure for new services • Participate in on-call and incident response, leading postmortems and reliability improvements

🎯 Requisitos

• 5+ years experience in SRE, DevOps, or infrastructure engineering roles • Strong Linux systems and networking fundamentals • Deep experience with Kubernetes and container orchestration at scale • Proficiency with Infrastructure-as-Code (Terraform, Helm, Ansible, etc.) • Strong automation and scripting skills (Python, Go, or Bash) • Experience with observability stacks (Prometheus, Grafana, Loki, Datadog, etc.) • Track record of operating production systems and leading incident response

🏖️ Benefícios

• Ownership and autonomy to shape systems • Opportunities to work directly with customers and providers

Candidatar-se

Vagas Similares

🕒 Fevereiro 25

WorkOS

51 - 200

🔌 API

🏢 Corporativo

🤝 B2B

Site Reliability Engineer ensuring reliability and performance at WorkOS across complex systems. Leading incident response and collaborating with cross-functional teams for operational excellence.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $175.000 - $275.000 / ano

💰 $80.000.000 Series B - WorkOS em 2022-05

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Fevereiro 25

Tactibit Technologies

11 - 50

💼 Consultoria

📦 Logística

🎖️ Defesa

DevOps Engineer working at Tactibit Technologies to modernize legacy architectures for mission-critical systems. Collaborate with teams on cloud migrations and automating business processes.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Fevereiro 17

Ensono

1001 - 5000

💼 Consultoria

DevOps Engineer working with AWS technologies on client deployment projects. Responsible for automation, support, and high availability of mission-critical solutions.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Fevereiro 17

Ensono

1001 - 5000

💼 Consultoria

Senior DevOps Engineer focused on deploying and supporting AWS technologies. Leading engineering projects and providing client support in a remote U.S. environment.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Fevereiro 17

Ensono

1001 - 5000

💼 Consultoria

AWS DevOps Engineer focused on deployment and automation of mission-critical solutions in AWS. Collaborating with teams for household name clients using the latest technologies.

🗣️🇺🇸🇬🇧 Inglês obrigatório