Staff HPC Engineer

🕒 vor 5 Tagen

🏄 California, Texas – Remote

infoinfo

⏰ Vollzeit

🔴 Experte

👷🏻‍♀️ Ingenieur

👻 Geisterscore 14%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 Mitarbeiter

💼 Beratung

📦 Logistik

🏗️ Bauwesen

💰 Post-IPO Equity im 2023-05

Consulting • Logistics • Construction

Bitdeer Technologies Group (Nasdaq: BTDR) ist ein führendes Unternehmen in der Blockchain- und Hochleistungs-Computerindustrie. Es ist einer der weltweit größten Inhaber von proprietärer Hash-Rate und Anbieter von Hash-Rate. Bitdeer hat sich der Bereitstellung umfassender Computerlösungen für seine Kunden verschrieben. Das Unternehmen wurde von Jihan Wu gegründet, einem frühen Befürworter und Pionier in der Kryptowährung, der mehrere führende Unternehmen mitbegründete, die die Blockchain-Wirtschaft bedienen. Matt Linghui Kong, der CEO der Bitdeer Group, führt das Unternehmen mit tiefem Branchenwissen und technologischem Fachwissen. Mit Hauptsitz in Singapur hat Bitdeer Mining-Datenzentren in den USA, Norwegen und Bhutan eingerichtet. Es bietet spezialisierte Mining-Infrastruktur, hochwertige Hash-Rate-Sharing-Produkte und zuverlässige Hosting-Dienste für globale Nutzer an. Das Unternehmen bietet außerdem fortschrittliche Cloud-Funktionen für Kunden mit hohen Anforderungen an künstliche Intelligenz. Engagement, Authentizität und Vertrauenswürdigkeit sind grundlegend für unsere Mission, der weltweit zuverlässigste Anbieter von umfassenden Blockchain- und Hochleistungs-Computing-Lösungen zu werden. Wir laden globales Talent ein, sich uns anzuschließen, um die Zukunft mitzugestalten.

Beschreibung

• Design, deploy, and operate production Slurm clusters on bare metal and virtual machines • Own Slurm cluster architecture, lifecycle, multi-tenant scheduling policy, and reliability across GPU infrastructure • Configure high availability, authentication, topology-aware GPU scheduling, accounting, partitions, QOS, fairshare, preemption, reservations, and tenant TRES limits • Lead adoption of the Slinky slurm-operator and evaluate slurm-bridge for Kubernetes workload co-scheduling • Shift GPU nodes between Slurm batch-training queues and Kubernetes inference capacity using elastic-capacity mechanisms • Operate Pyxis/Enroot and OCI/containerd job paths, supporting MPI/PMIx, module/Spack environments, and customer images • Build passive and active GPU-cluster health checks with automatic drain and job requeue • Own rack burn-in and acceptance testing before paid workloads are scheduled • Deliver reproducible clusters through Terraform, Ansible, golden images, PXE, Redfish, and IPMI • Instrument queue wait time, allocation efficiency, GPU utilization, and job failures through Prometheus/Grafana • Integrate Slurm accounting and GPU-hours with metering and invoicing pipelines • Write runbooks and tenant documentation, onboard and support enterprise customers, handle escalations, and mentor platform engineers

🎯 Anforderungen

• 8+ years in HPC, systems, or cloud infrastructure engineering • 4+ years operating production Slurm clusters at 100+ GPU-node scale with real users and service-level commitments • Deep hands-on expertise with slurm.conf, gres.conf, topology.conf, cgroup.conf, partitions, QOS, fairshare, preemption, reservations, slurmdbd accounting, slurmrestd, MUNGE/SACK and JWT authentication, and live-cluster version upgrades • Strong GPU and fabric fundamentals, including NVIDIA drivers, Fabric Manager, DCGM, MIG, InfiniBand/RoCEv2, subnet manager/UFM, rail-optimized topology, GPUDirect RDMA, and NCCL tuning and failure diagnosis • Production Kubernetes experience and working knowledge of operator/CRD patterns • Hands-on exposure to at least one Slurm-on-Kubernetes stack: Slinky slurm-operator or slurm-bridge, CoreWeave SUNK, or Nebius Soperator • Experience delivering bare-metal and virtualized compute, including provisioning, firmware/BIOS lifecycle management, KVM/QEMU or public-cloud-equivalent VM clusters, and Terraform/Ansible automation • Working knowledge of parallel and shared storage such as Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS • Proficient in Python and Bash for cluster automation • Clear written and verbal communication in English • Go experience is a plus

🏖️ Vorteile

• Equal employment opportunities in accordance with country, state, and local laws

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 5 Tagen

Hanwha Renewables

51 - 200

🏗️ Bauwesen

📦 Logistik

⚡ Energie

Principal Planning Engineer leading transmission studies, interconnection agreements, and due diligence for Hanwha Renewables’ utility-scale solar and storage projects. Supporting North American renewable energy development.

🇺🇸 Vereinigte Staaten – Remote

💵 $180.000 - $200.000 / Jahr

⏰ Vollzeit

🔴 Experte

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

RTOS

🕒 vor 5 Tagen

TerraPower

501 - 1000

⚡ Energie

💊 Pharmazie

Principal Nuclear Licensing Engineer leading U.S. NRC licensing for TerraPower’s Natrium advanced reactor. Driving regulatory strategy, licensing documents, and multidisciplinary approvals.

🇺🇸 Vereinigte Staaten – Remote

💵 $148.722 - $193.146 / Jahr

⏰ Vollzeit

🟠 Senior

🔴 Experte

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Tagen

GAI Consultants, Inc.

501 - 1000

💼 Beratung

📦 Logistik

🏭 Fertigung

Distribution Line Engineer designing overhead and underground distribution systems for GAI Consultants’ Arizona energy infrastructure projects. Supporting complex utility, generation, and industrial work from planning through delivery.

🇺🇸 Vereinigte Staaten – Remote

💰 Private equity im 2022-11

⏰ Vollzeit

🟠 Senior

🔴 Experte

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Tagen

RTX

10.000+ Mitarbeiter

🏭 Fertigung

💼 Beratung

📦 Logistik

Principal Mechanical Engineer supporting Patriot PAC-2 missile canister design, assembly, and testing. Representing Raytheon at supplier sites and resolving hardware-development challenges.

🇺🇸 Vereinigte Staaten – Remote

💵 $107.500 - $204.500 / Jahr

⏰ Vollzeit

🔴 Experte

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

Assembly

🕒 vor 6 Tagen

Cencora

10.000+ Mitarbeiter

💼 Beratung

📦 Logistik

🏥 Gesundheitswesen

Vulnerability management engineer advancing CTEM and reducing cyber exposure for Cencora, a global pharmaceutical solutions company. Analyzing enterprise vulnerabilities, driving remediation, and improving security posture across cloud, endpoint, network, application, and SaaS environments.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

🔴 Experte

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich