Staff HPC Engineer

🕒 Agosto 19

🏄 California, Texas – Remoto

infoinfo

⏰ Tempo Integral

🔴 Especialista

👷🏻‍♀️ Engenheiro

👻 Score fantasma 21%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 funcionários

💼 Consultoria

📦 Logística

🏗️ Construção

💰 Post-IPO Equity em 2023-05

Consulting • Logistics • Construction

A Bitdeer Technologies Group (Nasdaq: BTDR) é uma líder na indústria de blockchain e computação de alto desempenho. É uma das maiores detentoras mundiais de hash rate proprietário e fornecedoras de hash rate. A Bitdeer está comprometida em fornecer soluções de computação abrangentes para seus clientes.

Descrição

• Design, deploy, and operate production Slurm clusters on bare metal and virtual machines • Own Slurm cluster architecture, lifecycle, multi-tenant scheduling policy, and reliability across GPU infrastructure • Configure high availability, authentication, topology-aware GPU scheduling, accounting, partitions, QOS, fairshare, preemption, reservations, and tenant TRES limits • Lead adoption of the Slinky slurm-operator and evaluate slurm-bridge for Kubernetes workload co-scheduling • Shift GPU nodes between Slurm batch-training queues and Kubernetes inference capacity using elastic-capacity mechanisms • Operate Pyxis/Enroot and OCI/containerd job paths, supporting MPI/PMIx, module/Spack environments, and customer images • Build passive and active GPU-cluster health checks with automatic drain and job requeue • Own rack burn-in and acceptance testing before paid workloads are scheduled • Deliver reproducible clusters through Terraform, Ansible, golden images, PXE, Redfish, and IPMI • Instrument queue wait time, allocation efficiency, GPU utilization, and job failures through Prometheus/Grafana • Integrate Slurm accounting and GPU-hours with metering and invoicing pipelines • Write runbooks and tenant documentation, onboard and support enterprise customers, handle escalations, and mentor platform engineers

🎯 Requisitos

• 8+ years in HPC, systems, or cloud infrastructure engineering • 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments • Deep hands-on expertise with slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions, QOS, fairshare, preemption, reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and live-cluster version upgrades • Strong GPU and fabric fundamentals, including NVIDIA drivers, Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2, subnet manager/UFM, rail-optimized topology, GPUDirect RDMA, and NCCL tuning and failure diagnosis • Production Kubernetes experience and working knowledge of operator/CRD patterns • Hands-on exposure to at least one Slurm-on-Kubernetes stack: Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator • Experience delivering bare-metal and virtualized compute, including provisioning, firmware/BIOS lifecycle management, KVM/QEMU or public-cloud-equivalent VM clusters, and Terraform/Ansible automation • Working knowledge of parallel and shared storage such as Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS • Proficient in Python and Bash for cluster automation • Clear written and verbal communication in English • Go experience is a plus

🏖️ Benefícios

• Equal employment opportunities in accordance with country, state, and local laws

Candidatar-se

Vagas Similares

🕒 Agosto 19

Hanwha Renewables

51 - 200

🏗️ Construção

📦 Logística

⚡ Energia

Principal Planning Engineer leading transmission studies, interconnection agreements, and due diligence for Hanwha Renewables’ utility-scale solar and storage projects. Supporting North American renewable energy development.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $180.000 - $200.000 / ano

⏰ Tempo Integral

🔴 Especialista

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

RTOS

🕒 Agosto 19

TerraPower

501 - 1000

⚡ Energia

💊 Farmacêutico

Principal Nuclear Licensing Engineer advancing TerraPower’s Natrium reactor licensing strategy. Leading NRC applications, regulatory engagements, and licensing teams for nuclear technology company TerraPower.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $148.722 - $193.146 / ano

⏰ Tempo Integral

🔴 Especialista

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 19

GAI Consultants, Inc.

501 - 1000

💼 Consultoria

📦 Logística

🏭 Manufatura

Distribution Line Engineer designing overhead and underground distribution systems for GAI Consultants’ Arizona energy infrastructure projects. Supporting complex utility, generation, and industrial work from planning through delivery.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 Private equity em 2022-11

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 18

Cencora

10.000+ funcionários

💼 Consultoria

📦 Logística

🏥 Saúde

Engineer III strengthening Cencora’s healthcare-focused enterprise vulnerability and cyber exposure management. Leading CTEM, risk prioritization, remediation governance, and security reporting.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 18

CEQEL Critical Commissioning, LLC

11 - 50

💼 Consultoria

🏭 Manufatura

📦 Logística

Mechanical commissioning engineer leading full life cycle testing for CEQEL’s mission-critical facilities. Reviewing designs, managing schedules, and directing field commissioning with up to 75% domestic travel.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $110.000 - $140.000 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório