Forward Deployed Engineer – SRE

Vaga não está no LinkedIn

🕒 Julho 31

🏄 California – Remoto

info

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Andromeda

Andromeda

11 - 50 funcionários

🏥 Saúde

💼 Consultoria

🏨 Hospitalidade

🔥 Investimento no último ano

💰 $15.142.238 Series A - Andromeda Robotics em 2025-09

Healthcare • Consulting • Hospitality

A Andromeda é um serviço de computação com GPU e um marketplace que oferece acesso instantâneo a grandes clusters de aceleradores H100, H200 e B200 para experimentos, treinamentos em larga escala e inferência. Suporta orquestração com Slurm, Kubernetes ou SSH direto, oferece uso flexível sem duração mínima e preços competitivos, inclui expertise em DevOps, armazenamento local NAS ou transmitido sem taxas de entrada/saída, e suporte 24/7 com SLAs do setor. A empresa também opera um marketplace de GPUs de terceiros em gpulist. ai.

Descrição

• Serve as the primary technical point of contact for teams running large-scale training and inference workloads • Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns • Profile and improve distributed training performance on live workloads • Own reliability outcomes for the accounts you're deployed on • Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) • Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks • Turn every repeated deployment problem into automation

🎯 Requisitos

• Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent) • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training • Working knowledge of how large training and inference jobs actually run • Expert-level Linux experience: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at the syscall and hardware level • Strong experience running Kubernetes in production with GPU workloads • Strong engineering skills in Python, Go, or Bash • Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent) • Hands-on experience building monitoring and alerting for GPU-specific telemetry

🏖️ Benefícios

• Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage • 401(k) • Unlimited PTO

Candidatar-se

Vagas Similares

🕒 Julho 31

Smithfield Foods

10.000+ funcionários

🏭 Manufatura

🌾 Agricultura

🍽️ Alimentos e Bebidas

Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 31

CI&T

5001 - 10000

💼 Consultoria

🏥 Saúde

📣 Marketing

Lead design, deployment, and optimization of NVIDIA cuOpt solutions for AI-focused business operations. Collaborate across teams for effective routing and optimization solutions in Colombia.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $5.500.000 Venture Round em 2014-04

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 31

Rocket.net

11 - 50

☁️ SaaS

🛍️ Comércio Eletrônico

🏢 Corporativo

Site Reliability Engineer ensuring high standards for servers, services, and customer environments at Rocket.net. Providing advanced technical support and ensuring platform reliability.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 31

Branch

501 - 1000

💼 Consultoria

📣 Marketing

🔌 API

Senior Site Reliability Engineer at Branch improving platform reliability and scalability with automation. Collaborating with developers to enhance performance and observability of services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $175.000 - $185.000 / ano

💰 $282.000.000 Series F em 2022-02

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 31

Branch

501 - 1000

💼 Consultoria

📣 Marketing

🔌 API

Senior Database Reliability Engineer managing MySQL and CloudSQL systems. Ensuring high performance and availability in a remote role at Branch.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $175.000 - $185.000 / ano

💰 $282.000.000 Series F em 2022-02

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório