Cluster Engineer

Job not on LinkedIn

🕒 August 6

đŸ‡ș🇾 United States – Remote

⏰ Full Time

🟠 Senior

🔮 Lead

đŸ‘·đŸ»â€â™€ïž Engineer

đŸ‘» Ghost score 13%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of STN Incorporated

STN Incorporated

11 - 50 employees

Founded 2016

🏱 Enterprise

🔒 Cybersecurity

🔧 Hardware

Enterprise ‱ Cybersecurity ‱ Hardware

STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare.

📋 Description

‱ Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads ‱ Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency ‱ Optimize inference clusters for token generation throughput, low latency, and high GPU utilization ‱ Build and support production AI infrastructure running hundreds to thousands of GPUs ‱ Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers ‱ Perform NCCL benchmarking, analysis, and tuning for collective communication performance ‱ Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS ‱ Configure and tune distributed AI software stacks including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot ‱ Optimize GPU scheduling and resource allocation for training and inference environments ‱ Develop benchmarking and validation processes for hardware, firmware, drivers, and software releases ‱ Identify performance regressions and troubleshoot distributed training issues at scale ‱ Optimize storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O ‱ Work closely with ML engineers to improve training scalability and inference efficiency ‱ Create automation to deploy, validate, benchmark, and monitor GPU clusters ‱ Evaluate emerging AI infrastructure technologies and recommend platform architecture improvements

🎯 Requirements

‱ 7+ years designing or operating large-scale Linux infrastructure ‱ 5+ years supporting production GPU clusters for AI or HPC workloads ‱ Experience building multi-node GPU training environments from the ground up ‱ Deep expertise with distributed PyTorch training ‱ Extensive experience troubleshooting and optimizing NCCL communications ‱ Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications ‱ Experience benchmarking distributed training with nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf preferred ‱ Understanding of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism ‱ Experience optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth ‱ Experience tuning CUDA, NCCL, UCX, and MPI ‱ Expert-level Linux systems administration skills ‱ Experience with Slurm ‱ Experience using Pyxis and Enroot for containerized GPU workloads ‱ Strong Python and Bash scripting skills ‱ Strong understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet ‱ Experience designing or tuning AI storage, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance ‱ Experience with NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling ‱ Preferred: Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments

Apply Now

Similar Jobs

🕒 August 6

LangChain

11 - 50

đŸ€– Artificial Intelligence

đŸ€ B2B

☁ SaaS

Deployed Engineer partnering with LangChain customers to build, deploy, and operate production AI agents. Designing architectures, leading technical evaluations, and feeding field insights into LangChain’s platform.

đŸ‡ș🇾 United States – Remote

đŸ’” $150k - $250k / year

💰 $25M Series A on 2024-02

⏰ Full Time

🟡 Mid-level

🟠 Senior

đŸ‘·đŸ»â€â™€ïž Engineer

🕒 August 6

Kalles Group

11 - 50

🔒 Cybersecurity

🔐 Security

đŸ’Œ Consulting

Principal IAM Engineer Consultant leading secure, scalable CIAM design for regulated financial services clients. Defining authentication architectures, integration patterns, and engineering standards across delivery teams.

đŸ‡ș🇾 United States – Remote

đŸ’” $160k - $220k / year

⏰ Full Time

🔮 Lead

đŸ‘·đŸ»â€â™€ïž Engineer

🕒 August 5

OnCorps AI

51 - 200

💾 Finance

💳 Fintech

đŸ€– Artificial Intelligence

Solutions Engineer configuring AI-powered agents for fund operations at OnCorps, serving asset managers and fund administrators. Leading technical discovery, demonstrations, implementations, and API integrations for customer engagements.

đŸ‡ș🇾 United States – Remote

⏰ Full Time

🟠 Senior

🔮 Lead

đŸ‘·đŸ»â€â™€ïž Engineer

🕒 August 5

Autodesk

10,000+ employees

đŸ—ïž Construction

🏭 Manufacturing

đŸ’Œ Consulting

Distinguished Engineer leading Autodesk’s Product Access, entitlement, and licensing platforms for cloud and AI products. Shaping enterprise technical strategy, distributed systems, and AI-assisted engineering practices.

đŸ‡ș🇾 United States – Remote

đŸ’” $196k - $352.1k / year

⏰ Full Time

🟠 Senior

🔮 Lead

đŸ‘·đŸ»â€â™€ïž Engineer

🕒 August 5

Swarm Aero

51 - 200

🚀 Aerospace

đŸŽ–ïž Defense

🏭 Manufacturing

Forward-deployed engineer integrating and debugging C2 software for Swarm Aero’s autonomous swarming aircraft. Supporting U.S. government and military deployments across demanding domestic and international field environments.

đŸ‡ș🇾 United States – Remote

đŸ’” $180k - $240k / year

💰 Seed on 2024-09

⏰ Full Time

🟠 Senior

đŸ‘·đŸ»â€â™€ïž Engineer