Search Remote Jobs

Cluster Engineer

Job not on LinkedIn

đŸ”„ 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of STN Incorporated

STN Incorporated

11 - 50 employees

Founded 2016

🏱 Enterprise

🔒 Cybersecurity

🔧 Hardware

Enterprise ‱ Cybersecurity ‱ Hardware

STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare.

📋 Description

‱ Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads ‱ Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency ‱ Optimize inference clusters for token generation throughput, low latency, and high GPU utilization ‱ Build and support production AI infrastructure running hundreds to thousands of GPUs ‱ Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers ‱ Perform NCCL benchmarking, analysis, and tuning for collective communication performance ‱ Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS ‱ Configure and tune distributed AI software stacks including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot ‱ Optimize GPU scheduling and resource allocation for training and inference environments ‱ Develop benchmarking and validation processes for hardware, firmware, drivers, and software releases ‱ Identify performance regressions and troubleshoot distributed training issues at scale ‱ Optimize storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O ‱ Work closely with ML engineers to improve training scalability and inference efficiency ‱ Create automation to deploy, validate, benchmark, and monitor GPU clusters ‱ Evaluate emerging AI infrastructure technologies and recommend platform architecture improvements

🎯 Requirements

‱ 7+ years designing or operating large-scale Linux infrastructure ‱ 5+ years supporting production GPU clusters for AI or HPC workloads ‱ Experience building multi-node GPU training environments from the ground up ‱ Deep expertise with distributed PyTorch training ‱ Extensive experience troubleshooting and optimizing NCCL communications ‱ Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications ‱ Experience benchmarking distributed training with nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf preferred ‱ Understanding of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism ‱ Experience optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth ‱ Experience tuning CUDA, NCCL, UCX, and MPI ‱ Expert-level Linux systems administration skills ‱ Experience with Slurm ‱ Experience using Pyxis and Enroot for containerized GPU workloads ‱ Strong Python and Bash scripting skills ‱ Strong understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet ‱ Experience designing or tuning AI storage, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance ‱ Experience with NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling ‱ Preferred: Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments

Apply Now

Similar Jobs

đŸ”„ 29 minutes ago

LangChain

11 - 50

đŸ€– Artificial Intelligence

đŸ€ B2B

☁ SaaS

Deployed Engineer partnering with LangChain customers to build, deploy, and operate production AI agents. Designing architectures, leading technical evaluations, and feeding field insights into LangChain’s platform.

đŸ‡ș🇾 United States – Remote

đŸ’” $150k - $250k / year

💰 $25M Series A on 2024-02

⏰ Full Time

🟡 Mid-level

🟠 Senior

đŸ‘·đŸ»â€â™€ïž Engineer

đŸ”„ 2 hours ago

OnCorps AI

51 - 200

💾 Finance

💳 Fintech

đŸ€– Artificial Intelligence

Solutions Engineer configuring AI-powered agents for fund operations at OnCorps, serving asset managers and fund administrators. Leading technical discovery, demonstrations, implementations, and API integrations for customer engagements.

đŸ”„ 2 hours ago

Salas O'Brien

1001 - 5000

đŸ’Œ Consulting

đŸ—ïž Construction

đŸ„ Healthcare

Fire protection engineer managing designs for hyperscale data centers, federal, mission-critical, and industrial projects at Salas O’Brien. Coordinating technical deliverables, calculations, Revit/BIM modeling, and junior-engineer guidance remotely.

đŸ”„ 2 hours ago

Autodesk

10,000+ employees

đŸ—ïž Construction

🏭 Manufacturing

đŸ’Œ Consulting

Distinguished Engineer leading Autodesk’s Product Access, entitlement, and licensing platforms for cloud and AI products. Shaping enterprise technical strategy, distributed systems, and AI-assisted engineering practices.

đŸ”„ 3 hours ago

NBCUniversal

10,000+ employees

đŸ“± Media

Strategic Engineer troubleshooting enterprise network and voice issues for Comcast Business, a connectivity and managed-solutions provider. Analyzing OSI-layer faults, configuring call flows, and escalating complex incidents.