Tech Lead Manager, GPU Cluster Infrastructure

🕒 vor 18 Tagen

🏄 California – Remote

infoinfo

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

👻 Geisterscore 24%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of FAR.AI

FAR.AI

11 - 50 Mitarbeiter

Gegründet 2022

🤖 Künstliche Intelligenz

📚 Bildung

🤝 Non-Profit

Artificial Intelligence • Education • Non-profit

Ashby Electrical Limited ist ein 2018 gegründeter Elektrofachbetrieb, der sich auf kosteneffiziente und nachhaltige Elektrolösungen für Wohn- und Gewerbeprojekte spezialisiert hat. Das Unternehmen steht für hochwertige Arbeit über den gesamten Projektlebenszyklus hinweg – von der Planung über die Installation und Prüfung bis zur Inbetriebnahme. Ashby Electrical setzt auf starke, von Vertrauen, Integrität und Professionalität geprägte Kundenbeziehungen und stellt die termingerechte und budgetkonforme Umsetzung sicher.

Beschreibung

• Set the platform's technical direction and own its roadmap • Decide which systems to run, how to schedule and store across providers, and what to measure • Own architecture and scheduling and storage design • Debug failures across layers, including node health, GPU and fabric faults, and multi-node job hangs • Hire and grow a small team of senior engineers • Set priorities and ownership, scope projects, and provide regular feedback and coaching • Set security direction for a shared cluster where AI agents run experiments, including identity and access, workload isolation, and sandboxing • Define operations including on-call rotation, incident response, postmortems, fault tolerance, and observability • Participate in operational practices alongside the team • Serve as escalation point for research teams and infrastructure providers • Turn recurring problems into platform fixes • Work directly with researchers and engineers to keep large-scale experiments performant and fault-tolerant

🎯 Anforderungen

• Experience leading engineers as a manager, tech lead, or project lead, including setting technical direction, scoping work, and giving feedback • Desire to manage people directly; prior direct reports are not required • 5+ years of systems or infrastructure engineering experience on production Linux • Experience running GPU, HPC, or large-scale batch platforms • Experience owning systems from design through operation • Depth in at least one of scheduling, storage, networking, security, or GPU systems, with enough breadth to review designs in the others • Experience running production Kubernetes for GPU workloads with a batch layer such as Slurm, Kueue, or Volcano, including quotas, priority and preemption, and node health • Experience owning infrastructure as code and observability for a production fleet • Experience with Terraform or Ansible, Helm, ArgoCD, and Prometheus, or equivalents • Strong programming ability in Python, Go, Rust, C++, or another language commonly used for infrastructure • Clear technical writing for engineers, researchers, and providers • Additional desirable skills include distributed training infrastructure, distributed storage, cluster security, scheduler internals, multi-provider platforms, and greenfield team-building

🏖️ Vorteile

• Health insurance: 94% of insurance premium paid by the organization, commencing within 1 month after start date • 401(k) plan with up to 2% match • 25 days paid time off per year, accrued weekly • Up to 10 days of paid sick leave per year • Paid bereavement, family, medical, and pregnancy disability leave • Work computer and WFH stipend provided for eligible employees • Catered lunches and dinners on workdays at the Berkeley office • Visa sponsorship for in-person employees • Paid work trial lasting up to 1 week

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 18 Tagen

DTN

1001 - 5000

📦 Logistik

💼 Beratung

🏥 Gesundheitswesen

Senior Software Engineer building scalable Workday Finance services for DTN, a global operational data and technology company. Designing secure APIs, distributed systems, and financial workflows.

🇺🇸 Vereinigte Staaten – Remote

💵 $109.500 - $145.500 / Jahr

💰 Grant im 2014-09

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 18 Tagen

ServiceNow

10.000+ Mitarbeiter

🏢 Unternehmen

☁️ SaaS

🤖 Künstliche Intelligenz

Lead Technical Consultant designing and implementing ServiceNow solutions for client organizations. Leading delivery teams and driving projects to successful completion.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 18 Tagen

Lithic

51 - 200

💸 Finanzen

💳 Fintech

🔌 API

Senior Software Engineer building Lithic’s real-time card authorization and fraud-prevention systems. Developing scalable payment-risk tools, fraud controls, and distributed backend services.

🇺🇸 Vereinigte Staaten – Remote

💵 $160.000 - $200.000 / Jahr

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 18 Tagen

Cribl

501 - 1000

☁️ SaaS

Senior engineer building Cribl Stream integrations with Splunk, Kafka, and cloud storage. Developing NodeJS and TypeScript software for enterprise telemetry infrastructure.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 18 Tagen

Chainguard

51 - 200

🔐 Sicherheit

☁️ SaaS

🔒 Cybersecurity

Senior Software Engineer building AI-powered Go automation for Chainguard’s secure container image factory. Designing developer tooling, validation systems, and agentic pipelines for open-source supply chain security.

🗣️🇺🇸🇬🇧 Englisch erforderlich