
11 - 50 employees
🏥 Healthcare
💼 Consulting
🏨 Hospitality
💰 $15.1M Series A - Andromeda Robotics on 2025-09
Healthcare • Consulting • Hospitality
Andromeda is a GPU compute service and marketplace offering instant access to large clusters of H100, H200, and B200 accelerators for experiments, full-scale training, and inference. It supports orchestration with Slurm, Kubernetes, or direct SSH, provides flexible, no-minimum-duration usage and competitive pricing, and includes DevOps expertise, local NAS or streamed storage with no ingress/egress fees, and 24/7 support with industry SLAs. The company also operates a third-party GPU marketplace at gpulist. ai.
🕒 July 31
🏄 California – Remote
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
👻 Ghost score 10%
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
🏥 Healthcare
💼 Consulting
🏨 Hospitality
💰 $15.1M Series A - Andromeda Robotics on 2025-09
Healthcare • Consulting • Hospitality
Andromeda is a GPU compute service and marketplace offering instant access to large clusters of H100, H200, and B200 accelerators for experiments, full-scale training, and inference. It supports orchestration with Slurm, Kubernetes, or direct SSH, provides flexible, no-minimum-duration usage and competitive pricing, and includes DevOps expertise, local NAS or streamed storage with no ingress/egress fees, and 24/7 support with industry SLAs. The company also operates a third-party GPU marketplace at gpulist. ai.
• Serve as the primary technical point of contact for teams running large-scale training and inference workloads • Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns • Profile and improve distributed training performance on live workloads • Own reliability outcomes for the accounts you're deployed on • Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) • Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks • Turn every repeated deployment problem into automation
• Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent) • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training • Working knowledge of how large training and inference jobs actually run • Expert-level Linux experience: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at the syscall and hardware level • Strong experience running Kubernetes in production with GPU workloads • Strong engineering skills in Python, Go, or Bash • Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent) • Hands-on experience building monitoring and alerting for GPU-specific telemetry
• Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage • 401(k) • Unlimited PTO
Apply Now🕒 July 31
Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.
🕒 July 31
Site Reliability Engineer ensuring high standards for servers, services, and customer environments at Rocket.net. Providing advanced technical support and ensuring platform reliability.
🕒 July 31
DevOps Lead at Neural Earth responsible for maintaining CI/CD pipelines and monitoring AWS cloud infrastructure. Support incident response and operational work for smooth engineering processes.
🇺🇸 United States – Remote
💵 $125k - $156k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 July 31
Senior DevOps Engineer II at Bring a Trailer modernizing the infrastructure and security of a trusted automotive marketplace. Developing applications to enable engineering teams to work efficiently.
🇺🇸 United States – Remote
💵 $133k - $150k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 July 30
Senior Site Reliability Engineer at Pinterest ensuring reliability of cloud-native platforms on AWS and Kubernetes. Collaborating with teams to improve operational practices and incident responses.
🇺🇸 United States – Remote
💵 $139.8k - $287.7k / year
💰 Post IPO equity on 2022-08
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor