Cluster Engineer

Vaga não está no LinkedIn

🕒 Agosto 6

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

👷🏻‍♀️ Engenheiro

👻 Score fantasma 13%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of STN Incorporated

STN Incorporated

11 - 50 funcionários

Fundada em 2016

🏢 Corporativo

🔒 Cibersegurança

🔧 Hardware

Enterprise • Cybersecurity • Hardware

A STN Incorporated é um provedor de infraestrutura de TI gerenciada de nível empresarial e infraestrutura de nuvem que oferece infraestrutura segura e pronta para auditoria para sistemas críticos de negócios e workloads de IA exigentes. A STN opera um modelo operacional gerenciado oferecendo clouds privadas de CPU, infraestrutura GPU One AI, rede e armazenamento seguros, e suporte humano 24/7 com SLAs de alta disponibilidade. Seus serviços incluem infraestrutura gerenciada e operações de nuvem, operações de cibersegurança e resposta a incidentes, gestão de compliance e riscos (SOC 2 Tipo II, pronta para HIPAA), backup e recuperação, e aquisição de tecnologia empresarial e gestão de ciclo de vida. A STN atende a empresas, empresas SaaS de alto crescimento, construtores de IA e desenvolvedores de modelos, empresas de robótica/IA física e indústrias reguladas, como a área de saúde.

Descrição

• Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads • Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency • Optimize inference clusters for token generation throughput, low latency, and high GPU utilization • Build and support production AI infrastructure running hundreds to thousands of GPUs • Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers • Perform NCCL benchmarking, analysis, and tuning for collective communication performance • Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS • Configure and tune distributed AI software stacks including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot • Optimize GPU scheduling and resource allocation for training and inference environments • Develop benchmarking and validation processes for hardware, firmware, drivers, and software releases • Identify performance regressions and troubleshoot distributed training issues at scale • Optimize storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O • Work closely with ML engineers to improve training scalability and inference efficiency • Create automation to deploy, validate, benchmark, and monitor GPU clusters • Evaluate emerging AI infrastructure technologies and recommend platform architecture improvements

🎯 Requisitos

• 7+ years designing or operating large-scale Linux infrastructure • 5+ years supporting production GPU clusters for AI or HPC workloads • Experience building multi-node GPU training environments from the ground up • Deep expertise with distributed PyTorch training • Extensive experience troubleshooting and optimizing NCCL communications • Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications • Experience benchmarking distributed training with nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf preferred • Understanding of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism • Experience optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth • Experience tuning CUDA, NCCL, UCX, and MPI • Expert-level Linux systems administration skills • Experience with Slurm • Experience using Pyxis and Enroot for containerized GPU workloads • Strong Python and Bash scripting skills • Strong understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet • Experience designing or tuning AI storage, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance • Experience with NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling • Preferred: Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments

Candidatar-se

Vagas Similares

🕒 Agosto 6

LangChain

11 - 50

🤖 Inteligência Artificial

🤝 B2B

☁️ SaaS

Deployed Engineer partnering with LangChain customers to build, deploy, and operate production AI agents. Designing architectures, leading technical evaluations, and feeding field insights into LangChain’s platform.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $250.000 / ano

💰 $25.000.000 Series A em 2024-02

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 6

Kalles Group

11 - 50

🔒 Cibersegurança

🔐 Segurança

💼 Consultoria

Principal IAM Engineer Consultant leading secure, scalable CIAM design for regulated financial services clients. Defining authentication architectures, integration patterns, and engineering standards across delivery teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $160.000 - $220.000 / ano

⏰ Tempo Integral

🔴 Especialista

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 5

OnCorps AI

51 - 200

💸 Finanças

💳 Fintech

🤖 Inteligência Artificial

Solutions Engineer configuring AI-powered agents for fund operations at OnCorps, serving asset managers and fund administrators. Leading technical discovery, demonstrations, implementations, and API integrations for customer engagements.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 5

Swarm Aero

51 - 200

🚀 Aeroespacial

🎖️ Defesa

🏭 Manufatura

Forward-deployed engineer integrating and debugging C2 software for Swarm Aero’s autonomous swarming aircraft. Supporting U.S. government and military deployments across demanding domestic and international field environments.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $180.000 - $240.000 / ano

💰 Seed em 2024-09

⏰ Tempo Integral

🟠 Sênior

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 5

Jones Lang LaSalle Americas, Inc.

10.000+ funcionários

🏠 Imobiliário

🤝 B2B

💼 Consultoria

Senior Operating Engineer maintaining HVAC, mechanical, plumbing, and electrical systems for JLL’s Tennessee facilities clients. Managing preventive maintenance, troubleshooting, CMMS work orders, and energy-efficiency improvements.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $62.400 - $72.800 / ano

⏰ Tempo Integral

🟠 Sênior

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório