Tech Lead Manager, GPU Cluster Infrastructure

🕒 3 dias atrás

🏄 California – Remoto

infoinfo

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

👻 Score fantasma 24%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of FAR.AI

FAR.AI

11 - 50 funcionários

Fundada em 2022

🤖 Inteligência Artificial

📚 Educação

🤝 Sem Fins Lucrativos

Artificial Intelligence • Education • Non-profit

A Ashby Electrical Limited é uma empreiteira estabelecida em 2018 que se especializa em fornecer soluções elétricas econômicas e sustentáveis para projetos residenciais e comerciais. A empresa se orgulha de entregar trabalho de alta qualidade ao longo do ciclo de vida do projeto, incluindo design, instalação, teste e comissionamento. A Ashby Electrical está comprometida em construir relacionamentos sólidos com seus clientes, fundamentada na confiança, integridade e profissionalismo, garantindo que os projetos sejam concluídos dentro do prazo e do orçamento.

Descrição

• Set the platform's technical direction and own its roadmap • Decide which systems to run, how to schedule and store across providers, and what to measure • Own architecture and scheduling and storage design • Debug failures across layers, including node health, GPU and fabric faults, and multi-node job hangs • Hire and grow a small team of senior engineers • Set priorities and ownership, scope projects, and provide regular feedback and coaching • Set security direction for a shared cluster where AI agents run experiments, including identity and access, workload isolation, and sandboxing • Define operations including on-call rotation, incident response, postmortems, fault tolerance, and observability • Participate in operational practices alongside the team • Serve as escalation point for research teams and infrastructure providers • Turn recurring problems into platform fixes • Work directly with researchers and engineers to keep large-scale experiments performant and fault-tolerant

🎯 Requisitos

• Experience leading engineers as a manager, tech lead, or project lead, including setting technical direction, scoping work, and giving feedback • Desire to manage people directly; prior direct reports are not required • 5+ years of systems or infrastructure engineering experience on production Linux • Experience running GPU, HPC, or large-scale batch platforms • Experience owning systems from design through operation • Depth in at least one of scheduling, storage, networking, security, or GPU systems, with enough breadth to review designs in the others • Experience running production Kubernetes for GPU workloads with a batch layer such as Slurm, Kueue, or Volcano, including quotas, priority and preemption, and node health • Experience owning infrastructure as code and observability for a production fleet • Experience with Terraform or Ansible, Helm, ArgoCD, and Prometheus, or equivalents • Strong programming ability in Python, Go, Rust, C++, or another language commonly used for infrastructure • Clear technical writing for engineers, researchers, and providers • Additional desirable skills include distributed training infrastructure, distributed storage, cluster security, scheduler internals, multi-provider platforms, and greenfield team-building

🏖️ Benefícios

• Health insurance: 94% of insurance premium paid by the organization, commencing within 1 month after start date • 401(k) plan with up to 2% match • 25 days paid time off per year, accrued weekly • Up to 10 days of paid sick leave per year • Paid bereavement, family, medical, and pregnancy disability leave • Work computer and WFH stipend provided for eligible employees • Catered lunches and dinners on workdays at the Berkeley office • Visa sponsorship for in-person employees • Paid work trial lasting up to 1 week

Candidatar-se

Vagas Similares

🕒 3 dias atrás

DTN

1001 - 5000

📦 Logística

💼 Consultoria

🏥 Saúde

Senior Software Engineer building scalable Workday Finance services for DTN, a global operational data and technology company. Designing secure APIs, distributed systems, and financial workflows.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $109.500 - $145.500 / ano

💰 Grant em 2014-09

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

ServiceNow

10.000+ funcionários

🏢 Corporativo

☁️ SaaS

🤖 Inteligência Artificial

Lead Technical Consultant designing and implementing ServiceNow solutions for client organizations. Leading delivery teams and driving projects to successful completion.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Reddit, Inc.

501 - 1000

💼 Consultoria

📣 Marketing

📱 Mídia

Frontend engineer building Reddit products that help users start and grow communities. Developing scalable features across the full lifecycle with modern web technologies.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $190.800 - $267.100 / ano

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

GE Vernova

10.000+ funcionários

💼 Consultoria

📦 Logística

🏭 Manufatura

Technical Lead defining electrical, controls, and operability monitoring for GE Vernova’s wind turbine fleet. Guiding analytics, diagnostics, and reliability improvements across engineering teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $113.200 - $188.800 / ano

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Lithic

51 - 200

💸 Finanças

💳 Fintech

🔌 API

Senior Software Engineer building Lithic’s real-time card authorization and fraud-prevention systems. Developing scalable payment-risk tools, fraud controls, and distributed backend services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $160.000 - $200.000 / ano

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório