Search Remote Jobs

Staff HPC Engineer

Job not on LinkedIn

🔥 0 minutes ago

🏄 California, Texas – Remote

infoinfo

⏰ Full Time

đź”´ Lead

👷🏻‍♀️ Engineer

đź‘» Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 employees

đź’Ľ Consulting

📦 Logistics

🏗️ Construction

đź’° Post-IPO Equity on 2023-05

Consulting • Logistics • Construction

Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.

đź“‹ Description

• Design, deploy, and operate production Slurm clusters on bare metal and virtual machines • Own Slurm cluster architecture, lifecycle, multi-tenant scheduling policy, and reliability across GPU infrastructure • Configure high availability, authentication, topology-aware GPU scheduling, accounting, partitions, QOS, fairshare, preemption, reservations, and tenant TRES limits • Lead adoption of the Slinky slurm-operator and evaluate slurm-bridge for Kubernetes workload co-scheduling • Shift GPU nodes between Slurm batch-training queues and Kubernetes inference capacity using elastic-capacity mechanisms • Operate Pyxis/Enroot and OCI/containerd job paths, supporting MPI/PMIx, module/Spack environments, and customer images • Build passive and active GPU-cluster health checks with automatic drain and job requeue • Own rack burn-in and acceptance testing before paid workloads are scheduled • Deliver reproducible clusters through Terraform, Ansible, golden images, PXE, Redfish, and IPMI • Instrument queue wait time, allocation efficiency, GPU utilization, and job failures through Prometheus/Grafana • Integrate Slurm accounting and GPU-hours with metering and invoicing pipelines • Write runbooks and tenant documentation, onboard and support enterprise customers, handle escalations, and mentor platform engineers

🎯 Requirements

• 8+ years in HPC, systems, or cloud infrastructure engineering • 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments • Deep hands-on expertise with slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions, QOS, fairshare, preemption, reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and live-cluster version upgrades • Strong GPU and fabric fundamentals, including NVIDIA drivers, Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2, subnet manager/UFM, rail-optimized topology, GPUDirect RDMA, and NCCL tuning and failure diagnosis • Production Kubernetes experience and working knowledge of operator/CRD patterns • Hands-on exposure to at least one Slurm-on-Kubernetes stack: Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator • Experience delivering bare-metal and virtualized compute, including provisioning, firmware/BIOS lifecycle management, KVM/QEMU or public-cloud-equivalent VM clusters, and Terraform/Ansible automation • Working knowledge of parallel and shared storage such as Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS • Proficient in Python and Bash for cluster automation • Clear written and verbal communication in English • Go experience is a plus

🏖️ Benefits

• Equal employment opportunities in accordance with country, state, and local laws

Apply Now

Similar Jobs

🔥 2 hours ago

Hanwha Renewables

51 - 200

🏗️ Construction

📦 Logistics

⚡ Energy

Principal Planning Engineer leading transmission studies, interconnection agreements, and due diligence for Hanwha Renewables’ utility-scale solar and storage projects. Supporting North American renewable energy development.

🔥 4 hours ago

TerraPower

501 - 1000

⚡ Energy

đź’Š Pharmaceuticals

Principal Nuclear Licensing Engineer leading U.S. NRC licensing for TerraPower’s Natrium advanced reactor. Driving regulatory strategy, licensing documents, and multidisciplinary approvals.

🔥 7 hours ago

GAI Consultants, Inc.

501 - 1000

đź’Ľ Consulting

📦 Logistics

🏭 Manufacturing

Distribution Line Engineer designing overhead and underground distribution systems for GAI Consultants’ Arizona energy infrastructure projects. Supporting complex utility, generation, and industrial work from planning through delivery.

🔥 8 hours ago

RTX

10,000+ employees

🏭 Manufacturing

đź’Ľ Consulting

📦 Logistics

Principal Mechanical Engineer supporting Patriot PAC-2 missile canister design, assembly, and testing. Representing Raytheon at supplier sites and resolving hardware-development challenges.

đź•’ Yesterday

Cencora

10,000+ employees

đź’Ľ Consulting

📦 Logistics

🏥 Healthcare

Vulnerability management engineer advancing CTEM and reducing cyber exposure for Cencora, a global pharmaceutical solutions company. Analyzing enterprise vulnerabilities, driving remediation, and improving security posture across cloud, endpoint, network, application, and SaaS environments.