
11 - 50 funcionários
🏥 Saúde
💼 Consultoria
🏨 Hospitalidade
🔥 Investimento no último ano
💰 $15.142.238 Series A - Andromeda Robotics em 2025-09
Healthcare • Consulting • Hospitality
A Andromeda é um serviço de computação com GPU e um marketplace que oferece acesso instantâneo a grandes clusters de aceleradores H100, H200 e B200 para experimentos, treinamentos em larga escala e inferência. Suporta orquestração com Slurm, Kubernetes ou SSH direto, oferece uso flexível sem duração mínima e preços competitivos, inclui expertise em DevOps, armazenamento local NAS ou transmitido sem taxas de entrada/saída, e suporte 24/7 com SLAs do setor. A empresa também opera um marketplace de GPUs de terceiros em gpulist. ai.
🕒 Julho 31
🏄 California – Remoto
⏰ Tempo Integral
🟡 Pleno
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🦅 Patrocina Visto H1B
🗣️🇺🇸🇬🇧 Inglês obrigatório
Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

11 - 50 funcionários
🏥 Saúde
💼 Consultoria
🏨 Hospitalidade
🔥 Investimento no último ano
💰 $15.142.238 Series A - Andromeda Robotics em 2025-09
Healthcare • Consulting • Hospitality
A Andromeda é um serviço de computação com GPU e um marketplace que oferece acesso instantâneo a grandes clusters de aceleradores H100, H200 e B200 para experimentos, treinamentos em larga escala e inferência. Suporta orquestração com Slurm, Kubernetes ou SSH direto, oferece uso flexível sem duração mínima e preços competitivos, inclui expertise em DevOps, armazenamento local NAS ou transmitido sem taxas de entrada/saída, e suporte 24/7 com SLAs do setor. A empresa também opera um marketplace de GPUs de terceiros em gpulist. ai.
• Serve as the primary technical point of contact for teams running large-scale training and inference workloads • Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns • Profile and improve distributed training performance on live workloads • Own reliability outcomes for the accounts you're deployed on • Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) • Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks • Turn every repeated deployment problem into automation
• Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent) • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training • Working knowledge of how large training and inference jobs actually run • Expert-level Linux experience: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at the syscall and hardware level • Strong experience running Kubernetes in production with GPU workloads • Strong engineering skills in Python, Go, or Bash • Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent) • Hands-on experience building monitoring and alerting for GPU-specific telemetry
• Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage • 401(k) • Unlimited PTO
Candidatar-se🕒 Julho 31
Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.
🇺🇸 Estados Unidos – Remoto (EUA)
💵 $85.000 - $120.000 / ano
⏰ Tempo Integral
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🦅 Patrocina Visto H1B
🗣️🇺🇸🇬🇧 Inglês obrigatório
🕒 Julho 31
Lead design, deployment, and optimization of NVIDIA cuOpt solutions for AI-focused business operations. Collaborate across teams for effective routing and optimization solutions in Colombia.
🇺🇸 Estados Unidos – Remoto (EUA)
💰 $5.500.000 Venture Round em 2014-04
⏰ Tempo Integral
🟡 Pleno
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🗣️🇺🇸🇬🇧 Inglês obrigatório
🕒 Julho 31
Site Reliability Engineer ensuring high standards for servers, services, and customer environments at Rocket.net. Providing advanced technical support and ensuring platform reliability.
🇺🇸 Estados Unidos – Remoto (EUA)
⏰ Tempo Integral
🟡 Pleno
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🗣️🇺🇸🇬🇧 Inglês obrigatório
🕒 Julho 31
Senior Site Reliability Engineer at Branch improving platform reliability and scalability with automation. Collaborating with developers to enhance performance and observability of services.
🇺🇸 Estados Unidos – Remoto (EUA)
💵 $175.000 - $185.000 / ano
💰 $282.000.000 Series F em 2022-02
⏰ Tempo Integral
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🦅 Patrocina Visto H1B
🗣️🇺🇸🇬🇧 Inglês obrigatório
🕒 Julho 31
Senior Database Reliability Engineer managing MySQL and CloudSQL systems. Ensuring high performance and availability in a remote role at Branch.
🇺🇸 Estados Unidos – Remoto (EUA)
💵 $175.000 - $185.000 / ano
💰 $282.000.000 Series F em 2022-02
⏰ Tempo Integral
🟠 Sênior
⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)
🦅 Patrocina Visto H1B
🗣️🇺🇸🇬🇧 Inglês obrigatório