Senior Software Engineer, GPU Cluster Infrastructure

🕒 3 dias atrás

🌐 Estados Unidos, Singapura – Remoto

infoinfo

🏄 California – Remoto

infoinfo

💵 $150.000 - $275.000 / ano

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

👻 Score fantasma 14%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of FAR.AI

FAR.AI

11 - 50 funcionários

Fundada em 2022

🤖 Inteligência Artificial

📚 Educação

🤝 Sem Fins Lucrativos

Artificial Intelligence • Education • Non-profit

A Ashby Electrical Limited é uma empreiteira estabelecida em 2018 que se especializa em fornecer soluções elétricas econômicas e sustentáveis para projetos residenciais e comerciais. A empresa se orgulha de entregar trabalho de alta qualidade ao longo do ciclo de vida do projeto, incluindo design, instalação, teste e comissionamento. A Ashby Electrical está comprometida em construir relacionamentos sólidos com seus clientes, fundamentada na confiança, integridade e profissionalismo, garantindo que os projetos sejam concluídos dentro do prazo e do orçamento.

Descrição

• Operate the Kubernetes GPU fleet day to day, including node lifecycle, upgrades, driver and image rollouts, staged changes with safe rollback, and capacity planning • Own batch scheduling and multi-tenancy, including queues, quotas, priorities, preemption, gang scheduling, and fair share across research teams • Design and run storage under the fleet, including high-performance shared filesystems, object storage tiers, quotas, and backups • Keep multi-node training runs fault-tolerant by owning node health and automated draining, debugging NCCL and fabric problems, tracking stragglers and flaky GPUs, and building checkpoint and restart patterns • Harden the platform through identity and access, network policy, secrets, workload isolation, and sandboxing for AI agents • Bring new capacity online by acceptance-testing providers on fabric, NCCL, and storage throughput, holding them to SLAs, and integrating new clusters with infrastructure as code • Work directly with research teams on infrastructure problems and turn recurring issues into platform fixes • Share the on-call rotation, runbooks, and postmortems • Collaborate with researchers and engineers to keep large-scale experiments performant and fault-tolerant

🎯 Requisitos

• 3+ years in systems or infrastructure engineering on production Linux, running GPU, HPC, or large-scale batch platforms • Owned at least one system from design through operation • Production Kubernetes experience for GPU workloads with a batch layer such as Slurm, Kueue, Volcano, or similar • Experience with quotas, priority and preemption, and node health • Experience owning infrastructure as code and observability for a production fleet • Experience with Terraform or Ansible • Experience deploying with Helm and ArgoCD • Experience monitoring with Prometheus or equivalent tools • Strong programming skills in at least one infrastructure language such as Python, Go, Rust, or C++ • Ability to write clearly for engineers, researchers, and providers • Additional relevant expertise in distributed training infrastructure, distributed storage, cluster security, scheduler internals, or multi-provider platforms is advantageous

🏖️ Benefícios

• Health Insurance - 94% of Insurance premium paid by Organization commencing within 1 month after your start date • Retirement - 401(k) plan with up to 2% match • 25 days Paid Time Off per year, accrued weekly • Up to 10 days of paid sick leave per year • Paid Bereavement, Family, Medical and Pregnancy Disability Leave • WFH stipend and work computer provided for eligible employees • Catered lunches and dinners on workdays at the Berkeley office • Work-related travel and equipment expenses are covered • Visa sponsorship for in-person employees

Candidatar-se

Vagas Similares

🕒 3 dias atrás

DTN

1001 - 5000

📦 Logística

💼 Consultoria

🏥 Saúde

Senior Software Engineer building scalable Workday Finance services for DTN, a global operational data and technology company. Designing secure APIs, distributed systems, and financial workflows.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $109.500 - $145.500 / ano

💰 Grant em 2014-09

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

ServiceNow

10.000+ funcionários

🏢 Corporativo

☁️ SaaS

🤖 Inteligência Artificial

Lead Technical Consultant designing and implementing ServiceNow solutions for client organizations. Leading delivery teams and driving projects to successful completion.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Reddit, Inc.

501 - 1000

💼 Consultoria

📣 Marketing

📱 Mídia

Frontend engineer building Reddit products that help users start and grow communities. Developing scalable features across the full lifecycle with modern web technologies.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $190.800 - $267.100 / ano

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

GE Vernova

10.000+ funcionários

💼 Consultoria

📦 Logística

🏭 Manufatura

Technical Lead defining electrical, controls, and operability monitoring for GE Vernova’s wind turbine fleet. Guiding analytics, diagnostics, and reliability improvements across engineering teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $113.200 - $188.800 / ano

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Lithic

51 - 200

💸 Finanças

💳 Fintech

🔌 API

Senior Software Engineer building Lithic’s real-time card authorization and fraud-prevention systems. Developing scalable payment-risk tools, fraud controls, and distributed backend services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $160.000 - $200.000 / ano

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório