Search Remote Jobs

Tech Lead Manager, GPU Cluster Infrastructure

Job not on LinkedIn

🔥 1 minute ago

🏄 California – Remote

infoinfo

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 20%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of FAR.AI

FAR.AI

11 - 50 employees

Founded 2022

🤖 Artificial Intelligence

📚 Education

🤝 Non-profit

Artificial Intelligence • Education • Non-profit

FAR. AI is a research and education non-profit organization focused on ensuring advanced artificial intelligence is safe and beneficial for everyone. It conducts technical research on robustness, deception, interpretability, and evaluation of frontier AI systems, runs events and workshops that convene academic and industry leaders, and operates programs (like FAR. Labs and grantmaking) to support researchers and build capacity in trustworthy AI. FAR. AI also publishes findings, provides training and policy-facing convenings, and collaborates with universities and agencies to advance secure, aligned AI.

📋 Description

• Set the platform's technical direction and own its roadmap • Decide which systems to run, how to schedule and store across providers, and what to measure • Own architecture and scheduling and storage design • Debug failures across layers, including node health, GPU and fabric faults, and multi-node job hangs • Hire and grow a small team of senior engineers • Set priorities and ownership, scope projects, and provide regular feedback and coaching • Set security direction for a shared cluster where AI agents run experiments, including identity and access, workload isolation, and sandboxing • Define operations including on-call rotation, incident response, postmortems, fault tolerance, and observability • Participate in operational practices alongside the team • Serve as escalation point for research teams and infrastructure providers • Turn recurring problems into platform fixes • Work directly with researchers and engineers to keep large-scale experiments performant and fault-tolerant

🎯 Requirements

• Experience leading engineers as a manager, tech lead, or project lead, including setting technical direction, scoping work, and giving feedback • Desire to manage people directly; prior direct reports are not required • 5+ years of systems or infrastructure engineering experience on production Linux • Experience running GPU, HPC, or large-scale batch platforms • Experience owning systems from design through operation • Depth in at least one of scheduling, storage, networking, security, or GPU systems, with enough breadth to review designs in the others • Experience running production Kubernetes for GPU workloads with a batch layer such as Slurm, Kueue, or Volcano, including quotas, priority and preemption, and node health • Experience owning infrastructure as code and observability for a production fleet • Experience with Terraform or Ansible, Helm, ArgoCD, and Prometheus, or equivalents • Strong programming ability in Python, Go, Rust, C++, or another language commonly used for infrastructure • Clear technical writing for engineers, researchers, and providers • Additional desirable skills include distributed training infrastructure, distributed storage, cluster security, scheduler internals, multi-provider platforms, and greenfield team-building

🏖️ Benefits

• Health insurance: 94% of insurance premium paid by the organization, commencing within 1 month after start date • 401(k) plan with up to 2% match • 25 days paid time off per year, accrued weekly • Up to 10 days of paid sick leave per year • Paid bereavement, family, medical, and pregnancy disability leave • Work computer and WFH stipend provided for eligible employees • Catered lunches and dinners on workdays at the Berkeley office • Visa sponsorship for in-person employees • Paid work trial lasting up to 1 week

Apply Now

Similar Jobs

🔥 59 minutes ago

DTN

1001 - 5000

📦 Logistics

💼 Consulting

🏥 Healthcare

Senior Software Engineer building scalable Workday Finance services for DTN, a global operational data and technology company. Designing secure APIs, distributed systems, and financial workflows.

🔥 1 hour ago

ServiceNow

10,000+ employees

🏢 Enterprise

☁️ SaaS

🤖 Artificial Intelligence

Lead Technical Consultant designing and implementing ServiceNow solutions for client organizations. Leading delivery teams and driving projects to successful completion.

🔥 2 hours ago

Reddit, Inc.

501 - 1000

💼 Consulting

📣 Marketing

📱 Media

Frontend engineer building Reddit products that help users start and grow communities. Developing scalable features across the full lifecycle with modern web technologies.

🔥 2 hours ago

GE Vernova

10,000+ employees

💼 Consulting

📦 Logistics

🏭 Manufacturing

Technical Lead defining electrical, controls, and operability monitoring for GE Vernova’s wind turbine fleet. Guiding analytics, diagnostics, and reliability improvements across engineering teams.

🔥 2 hours ago

Lithic

51 - 200

💸 Finance

💳 Fintech

🔌 API

Senior Software Engineer building Lithic’s real-time card authorization and fraud-prevention systems. Developing scalable payment-risk tools, fraud controls, and distributed backend services.