Staff HPC Engineer

đź•’ August 19

🏄 California, Texas – Remote

infoinfo

⏰ Full Time

đź”´ Lead

👷🏻‍♀️ Engineer

đź‘» Ghost score 15%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 employees

đź’Ľ Consulting

📦 Logistics

🏗️ Construction

đź’° Post-IPO Equity on 2023-05

Consulting • Logistics • Construction

Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.

đź“‹ Description

• Design, deploy, and operate production Slurm clusters on bare metal and virtual machines • Own Slurm cluster architecture, lifecycle, multi-tenant scheduling policy, and reliability across GPU infrastructure • Configure high availability, authentication, topology-aware GPU scheduling, accounting, partitions, QOS, fairshare, preemption, reservations, and tenant TRES limits • Lead adoption of the Slinky slurm-operator and evaluate slurm-bridge for Kubernetes workload co-scheduling • Shift GPU nodes between Slurm batch-training queues and Kubernetes inference capacity using elastic-capacity mechanisms • Operate Pyxis/Enroot and OCI/containerd job paths, supporting MPI/PMIx, module/Spack environments, and customer images • Build passive and active GPU-cluster health checks with automatic drain and job requeue • Own rack burn-in and acceptance testing before paid workloads are scheduled • Deliver reproducible clusters through Terraform, Ansible, golden images, PXE, Redfish, and IPMI • Instrument queue wait time, allocation efficiency, GPU utilization, and job failures through Prometheus/Grafana • Integrate Slurm accounting and GPU-hours with metering and invoicing pipelines • Write runbooks and tenant documentation, onboard and support enterprise customers, handle escalations, and mentor platform engineers

🎯 Requirements

• 8+ years in HPC, systems, or cloud infrastructure engineering • 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments • Deep hands-on expertise with slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions, QOS, fairshare, preemption, reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and live-cluster version upgrades • Strong GPU and fabric fundamentals, including NVIDIA drivers, Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2, subnet manager/UFM, rail-optimized topology, GPUDirect RDMA, and NCCL tuning and failure diagnosis • Production Kubernetes experience and working knowledge of operator/CRD patterns • Hands-on exposure to at least one Slurm-on-Kubernetes stack: Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator • Experience delivering bare-metal and virtualized compute, including provisioning, firmware/BIOS lifecycle management, KVM/QEMU or public-cloud-equivalent VM clusters, and Terraform/Ansible automation • Working knowledge of parallel and shared storage such as Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS • Proficient in Python and Bash for cluster automation • Clear written and verbal communication in English • Go experience is a plus

🏖️ Benefits

• Equal employment opportunities in accordance with country, state, and local laws

Apply Now

Similar Jobs

đź•’ August 19

Hanwha Renewables

51 - 200

🏗️ Construction

📦 Logistics

⚡ Energy

Principal Planning Engineer leading transmission studies, interconnection agreements, and due diligence for Hanwha Renewables’ utility-scale solar and storage projects. Supporting North American renewable energy development.

🇺🇸 United States – Remote

đź’µ $180k - $200k / year

⏰ Full Time

đź”´ Lead

👷🏻‍♀️ Engineer

RTOS

đź•’ August 19

GAI Consultants, Inc.

501 - 1000

đź’Ľ Consulting

📦 Logistics

🏭 Manufacturing

Distribution Line Engineer designing overhead and underground distribution systems for GAI Consultants’ Arizona energy infrastructure projects. Supporting complex utility, generation, and industrial work from planning through delivery.

🇺🇸 United States – Remote

đź’° Private equity on 2022-11

⏰ Full Time

đźź  Senior

đź”´ Lead

👷🏻‍♀️ Engineer

đź•’ August 19

RTX

10,000+ employees

🏭 Manufacturing

đź’Ľ Consulting

📦 Logistics

Principal Mechanical Engineer supporting Patriot PAC-2 missile canister design, assembly, and testing. Representing Raytheon at supplier sites and resolving hardware-development challenges.

🇺🇸 United States – Remote

đź’µ $107.5k - $204.5k / year

⏰ Full Time

đź”´ Lead

👷🏻‍♀️ Engineer

Assembly

đź•’ August 18

Cencora

10,000+ employees

đź’Ľ Consulting

📦 Logistics

🏥 Healthcare

Vulnerability management engineer advancing CTEM and reducing cyber exposure for Cencora, a global pharmaceutical solutions company. Analyzing enterprise vulnerabilities, driving remediation, and improving security posture across cloud, endpoint, network, application, and SaaS environments.

🇺🇸 United States – Remote

⏰ Full Time

đźź  Senior

đź”´ Lead

👷🏻‍♀️ Engineer

đź•’ August 18

CEQEL Critical Commissioning, LLC

11 - 50

đź’Ľ Consulting

🏭 Manufacturing

📦 Logistics

Mechanical commissioning engineer leading full life cycle testing for CEQEL’s mission-critical facilities. Reviewing designs, managing schedules, and directing field commissioning with up to 75% domestic travel.

🇺🇸 United States – Remote

đź’µ $110k - $140k / year

⏰ Full Time

đźź  Senior

đź”´ Lead

👷🏻‍♀️ Engineer