Senior Software Engineer, GPU Cluster Infrastructure

Job not on LinkedIn

🔥 1 hour ago

🌐 United States, Singapore – Remote

infoinfo

🏄 California – Remote

infoinfo

💵 $150k - $275k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of FAR.AI

FAR.AI

11 - 50 employees

Founded 2022

🤖 Artificial Intelligence

📚 Education

🤝 Non-profit

Artificial Intelligence • Education • Non-profit

FAR. AI is a research and education non-profit organization focused on ensuring advanced artificial intelligence is safe and beneficial for everyone. It conducts technical research on robustness, deception, interpretability, and evaluation of frontier AI systems, runs events and workshops that convene academic and industry leaders, and operates programs (like FAR. Labs and grantmaking) to support researchers and build capacity in trustworthy AI. FAR. AI also publishes findings, provides training and policy-facing convenings, and collaborates with universities and agencies to advance secure, aligned AI.

📋 Description

• Operate the Kubernetes GPU fleet day to day, including node lifecycle, upgrades, driver and image rollouts, staged changes with safe rollback, and capacity planning • Own batch scheduling and multi-tenancy, including queues, quotas, priorities, preemption, gang scheduling, and fair share across research teams • Design and run storage under the fleet, including high-performance shared filesystems, object storage tiers, quotas, and backups • Keep multi-node training runs fault-tolerant by owning node health and automated draining, debugging NCCL and fabric problems, tracking stragglers and flaky GPUs, and building checkpoint and restart patterns • Harden the platform through identity and access, network policy, secrets, workload isolation, and sandboxing for AI agents • Bring new capacity online by acceptance-testing providers on fabric, NCCL, and storage throughput, holding them to SLAs, and integrating new clusters with infrastructure as code • Work directly with research teams on infrastructure problems and turn recurring issues into platform fixes • Share the on-call rotation, runbooks, and postmortems • Collaborate with researchers and engineers to keep large-scale experiments performant and fault-tolerant

🎯 Requirements

• 3+ years in systems or infrastructure engineering on production Linux, running GPU, HPC, or large-scale batch platforms • Owned at least one system from design through operation • Production Kubernetes experience for GPU workloads with a batch layer such as Slurm, Kueue, Volcano, or similar • Experience with quotas, priority and preemption, and node health • Experience owning infrastructure as code and observability for a production fleet • Experience with Terraform or Ansible • Experience deploying with Helm and ArgoCD • Experience monitoring with Prometheus or equivalent tools • Strong programming skills in at least one infrastructure language such as Python, Go, Rust, or C++ • Ability to write clearly for engineers, researchers, and providers • Additional relevant expertise in distributed training infrastructure, distributed storage, cluster security, scheduler internals, or multi-provider platforms is advantageous

🏖️ Benefits

• Health Insurance - 94% of Insurance premium paid by Organization commencing within 1 month after your start date • Retirement - 401(k) plan with up to 2% match • 25 days Paid Time Off per year, accrued weekly • Up to 10 days of paid sick leave per year • Paid Bereavement, Family, Medical and Pregnancy Disability Leave • WFH stipend and work computer provided for eligible employees • Catered lunches and dinners on workdays at the Berkeley office • Work-related travel and equipment expenses are covered • Visa sponsorship for in-person employees

Apply Now

Similar Jobs

🔥 2 hours ago

DTN

1001 - 5000

📦 Logistics

💼 Consulting

🏥 Healthcare

Senior Software Engineer building scalable Workday Finance services for DTN, a global operational data and technology company. Designing secure APIs, distributed systems, and financial workflows.

🔥 3 hours ago

ServiceNow

10,000+ employees

🏢 Enterprise

☁️ SaaS

🤖 Artificial Intelligence

Lead Technical Consultant designing and implementing ServiceNow solutions for client organizations. Leading delivery teams and driving projects to successful completion.

🔥 3 hours ago

Reddit, Inc.

501 - 1000

💼 Consulting

📣 Marketing

📱 Media

Frontend engineer building Reddit products that help users start and grow communities. Developing scalable features across the full lifecycle with modern web technologies.

🔥 3 hours ago

GE Vernova

10,000+ employees

💼 Consulting

📦 Logistics

🏭 Manufacturing

Technical Lead defining electrical, controls, and operability monitoring for GE Vernova’s wind turbine fleet. Guiding analytics, diagnostics, and reliability improvements across engineering teams.

🔥 4 hours ago

Lithic

51 - 200

💸 Finance

💳 Fintech

🔌 API

Senior Software Engineer building Lithic’s real-time card authorization and fraud-prevention systems. Developing scalable payment-risk tools, fraud controls, and distributed backend services.