
11 - 50 employees
🏥 Healthcare
💼 Consulting
🏨 Hospitality
🔥 Funding within the last year
💰 $15.1M Series A - Andromeda Robotics on 2025-09
Healthcare • Consulting • Hospitality
Andromeda is a GPU compute service and marketplace offering instant access to large clusters of H100, H200, and B200 accelerators for experiments, full-scale training, and inference. It supports orchestration with Slurm, Kubernetes, or direct SSH, provides flexible, no-minimum-duration usage and competitive pricing, and includes DevOps expertise, local NAS or streamed storage with no ingress/egress fees, and 24/7 support with industry SLAs. The company also operates a third-party GPU marketplace at gpulist. ai.
🔥 1 hour ago
🏄 California – Remote
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
🏥 Healthcare
💼 Consulting
🏨 Hospitality
🔥 Funding within the last year
💰 $15.1M Series A - Andromeda Robotics on 2025-09
Healthcare • Consulting • Hospitality
Andromeda is a GPU compute service and marketplace offering instant access to large clusters of H100, H200, and B200 accelerators for experiments, full-scale training, and inference. It supports orchestration with Slurm, Kubernetes, or direct SSH, provides flexible, no-minimum-duration usage and competitive pricing, and includes DevOps expertise, local NAS or streamed storage with no ingress/egress fees, and 24/7 support with industry SLAs. The company also operates a third-party GPU marketplace at gpulist. ai.
• Serve as the primary technical point of contact for teams running large-scale training and inference workloads • Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns • Profile and improve distributed training performance on live workloads • Own reliability outcomes for the accounts you're deployed on • Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) • Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks • Turn every repeated deployment problem into automation
• Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent) • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training • Working knowledge of how large training and inference jobs actually run • Expert-level Linux experience: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at the syscall and hardware level • Strong experience running Kubernetes in production with GPU workloads • Strong engineering skills in Python, Go, or Bash • Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent) • Hands-on experience building monitoring and alerting for GPU-specific telemetry
• Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage • 401(k) • Unlimited PTO
Apply Now🔥 3 hours ago
AWS DevOps Engineer responsible for designing and maintaining AWS cloud environments for Marine Corps IT systems. Collaborating across teams to ensure performance, security, and compliance needs are met.
🇺🇸 United States – Remote
⏰ Full Time
🟠 Senior
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
AWS
Cloud
Docker
Kubernetes
Python
Terraform
🔥 6 hours ago
Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.
🔥 6 hours ago
Lead design, deployment, and optimization of NVIDIA cuOpt solutions for AI-focused business operations. Collaborate across teams for effective routing and optimization solutions in Colombia.
🇺🇸 United States – Remote
💰 $5.5M Venture Round on 2014-04
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Azure
Cloud
🔥 7 hours ago
Site Reliability Engineer ensuring high standards for servers, services, and customer environments at Rocket.net. Providing advanced technical support and ensuring platform reliability.
Apache
DNS
Linux
MariaDB
MySQL
NGINX
PHP
Redis
WordPress
🔥 7 hours ago
Senior Site Reliability Engineer at Branch improving platform reliability and scalability with automation. Collaborating with developers to enhance performance and observability of services.
🇺🇸 United States – Remote
💵 $175k - $185k / year
💰 $282M Series F on 2022-02
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Docker
Gradle
Java
Kubernetes
Spring
Spring Boot
SpringBoot
Terraform
Go