Forward Deployed Engineer – SRE

Job not on LinkedIn

🔥 0 minutes ago

🏄 California – Remote

info

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Andromeda

Andromeda

11 - 50 employees

🏥 Healthcare

💼 Consulting

🏨 Hospitality

🔥 Funding within the last year

💰 $15.1M Series A - Andromeda Robotics on 2025-09

Healthcare • Consulting • Hospitality

Andromeda is a GPU compute service and marketplace offering instant access to large clusters of H100, H200, and B200 accelerators for experiments, full-scale training, and inference. It supports orchestration with Slurm, Kubernetes, or direct SSH, provides flexible, no-minimum-duration usage and competitive pricing, and includes DevOps expertise, local NAS or streamed storage with no ingress/egress fees, and 24/7 support with industry SLAs. The company also operates a third-party GPU marketplace at gpulist. ai.

📋 Description

• Serve as the primary technical point of contact for teams running large-scale training and inference workloads • Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns • Profile and improve distributed training performance on live workloads • Own reliability outcomes for the accounts you're deployed on • Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) • Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks • Turn every repeated deployment problem into automation

🎯 Requirements

• Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent) • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training • Working knowledge of how large training and inference jobs actually run • Expert-level Linux experience: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at the syscall and hardware level • Strong experience running Kubernetes in production with GPU workloads • Strong engineering skills in Python, Go, or Bash • Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent) • Hands-on experience building monitoring and alerting for GPU-specific telemetry

🏖️ Benefits

• Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage • 401(k) • Unlimited PTO

Apply Now

Similar Jobs

🔥 1 hour ago

Millennium

201 - 500

💼 Consulting

🎖️ Defense

🔒 Cybersecurity

AWS DevOps Engineer responsible for designing and maintaining AWS cloud environments for Marine Corps IT systems. Collaborating across teams to ensure performance, security, and compliance needs are met.

🔥 4 hours ago

Smithfield Foods

10,000+ employees

🏭 Manufacturing

🌾 Agriculture

🍽️ Food & Beverage

Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.

🔥 4 hours ago

CI&T

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Lead design, deployment, and optimization of NVIDIA cuOpt solutions for AI-focused business operations. Collaborate across teams for effective routing and optimization solutions in Colombia.

🔥 6 hours ago

Rocket.net

11 - 50

☁️ SaaS

🛍️ eCommerce

🏢 Enterprise

Site Reliability Engineer ensuring high standards for servers, services, and customer environments at Rocket.net. Providing advanced technical support and ensuring platform reliability.

🔥 6 hours ago

Branch

501 - 1000

💼 Consulting

📣 Marketing

🔌 API

Senior Site Reliability Engineer at Branch improving platform reliability and scalability with automation. Collaborating with developers to enhance performance and observability of services.