Cluster Engineer

Stelle nicht auf LinkedIn

🕒 vor 1 Tag

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

🔴 Experte

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of STN Incorporated

STN Incorporated

11 - 50 Mitarbeiter

Gegründet 2016

🏢 Unternehmen

🔒 Cybersecurity

🔧 Hardware

Enterprise • Cybersecurity • Hardware

STN Incorporated ist ein Anbieter für IT- und Cloud-Infrastruktur auf Unternehmensniveau, der sichere, prüfbereite Infrastrukturen für geschäftskritische Systeme und anspruchsvolle KI-Arbeitslasten liefert. STN betreibt ein Managed-Operating-Modell, das private CPU-Clouds, GPU One AI-Infrastruktur, sichere Netzwerke und Speicher sowie 24/7 menschlichen Support mit hohen Verfügbarkeits-SLAs bietet. Zu den Dienstleistungen gehören verwaltete Infrastruktur- und Cloud-Operationen, Cybersecurity-Operationen und Vorfallreaktionen, Compliance- und Risikomanagement (SOC 2 Typ II, HIPAA-konform), Backup und Wiederherstellung sowie Beschaffung und Lebenszyklusmanagement von Unternehmens-Technologien. STN bedient Unternehmen, wachstumsstarke SaaS-Unternehmen, KI-Entwickler und Modellbauer, Robotik-/physische KI-Firmen und regulierte Branchen wie das Gesundheitswesen.

Beschreibung

• Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads • Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency • Optimize inference clusters for token generation throughput, low latency, and high GPU utilization • Build and support production AI infrastructure running hundreds to thousands of GPUs • Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers • Perform NCCL benchmarking, analysis, and tuning for collective communication performance • Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS • Configure and tune distributed AI software stacks including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot • Optimize GPU scheduling and resource allocation for training and inference environments • Develop benchmarking and validation processes for hardware, firmware, drivers, and software releases • Identify performance regressions and troubleshoot distributed training issues at scale • Optimize storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O • Work closely with ML engineers to improve training scalability and inference efficiency • Create automation to deploy, validate, benchmark, and monitor GPU clusters • Evaluate emerging AI infrastructure technologies and recommend platform architecture improvements

🎯 Anforderungen

• 7+ years designing or operating large-scale Linux infrastructure • 5+ years supporting production GPU clusters for AI or HPC workloads • Experience building multi-node GPU training environments from the ground up • Deep expertise with distributed PyTorch training • Extensive experience troubleshooting and optimizing NCCL communications • Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications • Experience benchmarking distributed training with nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf preferred • Understanding of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism • Experience optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth • Experience tuning CUDA, NCCL, UCX, and MPI • Expert-level Linux systems administration skills • Experience with Slurm • Experience using Pyxis and Enroot for containerized GPU workloads • Strong Python and Bash scripting skills • Strong understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet • Experience designing or tuning AI storage, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance • Experience with NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling • Preferred: Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Tag

LangChain

11 - 50

🤖 Künstliche Intelligenz

🤝 B2B

☁️ SaaS

Deployed Engineer partnering with LangChain customers to build, deploy, and operate production AI agents. Designing architectures, leading technical evaluations, and feeding field insights into LangChain’s platform.

🇺🇸 Vereinigte Staaten – Remote

💵 $150.000 - $250.000 / Jahr

💰 €25.000.000 Series A im 2024-02

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Tag

Kalles Group

11 - 50

🔒 Cybersecurity

🔐 Sicherheit

💼 Beratung

Principal IAM Engineer Consultant leading secure, scalable CIAM design for regulated financial services clients. Defining authentication architectures, integration patterns, and engineering standards across delivery teams.

🇺🇸 Vereinigte Staaten – Remote

💵 $160.000 - $220.000 / Jahr

⏰ Vollzeit

🔴 Experte

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

Azure

Cloud

Cyber Security

🕒 vor 1 Tag

OnCorps AI

51 - 200

💸 Finanzen

💳 Fintech

🤖 Künstliche Intelligenz

Solutions Engineer configuring AI-powered agents for fund operations at OnCorps, serving asset managers and fund administrators. Leading technical discovery, demonstrations, implementations, and API integrations for customer engagements.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

🔴 Experte

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Tag

Salas O'Brien

1001 - 5000

💼 Beratung

🏗️ Bauwesen

🏥 Gesundheitswesen

Fire protection engineer managing designs for hyperscale data centers, federal, mission-critical, and industrial projects at Salas O’Brien. Coordinating technical deliverables, calculations, Revit/BIM modeling, and junior-engineer guidance remotely.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Tag

Autodesk

10.000+ Mitarbeiter

🏗️ Bauwesen

🏭 Fertigung

💼 Beratung

Distinguished Engineer leading Autodesk’s Product Access, entitlement, and licensing platforms for cloud and AI products. Shaping enterprise technical strategy, distributed systems, and AI-assisted engineering practices.

🇺🇸 Vereinigte Staaten – Remote

💵 $196.000 - $352.110 / Jahr

⏰ Vollzeit

🟠 Senior

🔴 Experte

👷🏻‍♀️ Ingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich