
11 - 50 employees
Founded 2016
đą Enterprise
đ Cybersecurity
đ§ Hardware
Enterprise âą Cybersecurity âą Hardware
STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare.
đ„ 0 minutes ago
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
Founded 2016
đą Enterprise
đ Cybersecurity
đ§ Hardware
Enterprise âą Cybersecurity âą Hardware
STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare.
âą Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads âą Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency âą Optimize inference clusters for token generation throughput, low latency, and high GPU utilization âą Build and support production AI infrastructure running hundreds to thousands of GPUs âą Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers âą Perform NCCL benchmarking, analysis, and tuning for collective communication performance âą Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS âą Configure and tune distributed AI software stacks including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot âą Optimize GPU scheduling and resource allocation for training and inference environments âą Develop benchmarking and validation processes for hardware, firmware, drivers, and software releases âą Identify performance regressions and troubleshoot distributed training issues at scale âą Optimize storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O âą Work closely with ML engineers to improve training scalability and inference efficiency âą Create automation to deploy, validate, benchmark, and monitor GPU clusters âą Evaluate emerging AI infrastructure technologies and recommend platform architecture improvements
âą 7+ years designing or operating large-scale Linux infrastructure âą 5+ years supporting production GPU clusters for AI or HPC workloads âą Experience building multi-node GPU training environments from the ground up âą Deep expertise with distributed PyTorch training âą Extensive experience troubleshooting and optimizing NCCL communications âą Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications âą Experience benchmarking distributed training with nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf preferred âą Understanding of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism âą Experience optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth âą Experience tuning CUDA, NCCL, UCX, and MPI âą Expert-level Linux systems administration skills âą Experience with Slurm âą Experience using Pyxis and Enroot for containerized GPU workloads âą Strong Python and Bash scripting skills âą Strong understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet âą Experience designing or tuning AI storage, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance âą Experience with NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling âą Preferred: Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments
Apply Nowđ„ 29 minutes ago
Deployed Engineer partnering with LangChain customers to build, deploy, and operate production AI agents. Designing architectures, leading technical evaluations, and feeding field insights into LangChainâs platform.
đșđž United States â Remote
đ” $150k - $250k / year
đ° $25M Series A on 2024-02
â° Full Time
đĄ Mid-level
đ Senior
đ·đ»ââïž Engineer
đ„ 2 hours ago
Solutions Engineer configuring AI-powered agents for fund operations at OnCorps, serving asset managers and fund administrators. Leading technical discovery, demonstrations, implementations, and API integrations for customer engagements.
đ„ 2 hours ago
Fire protection engineer managing designs for hyperscale data centers, federal, mission-critical, and industrial projects at Salas OâBrien. Coordinating technical deliverables, calculations, Revit/BIM modeling, and junior-engineer guidance remotely.
đ„ 2 hours ago
Distinguished Engineer leading Autodeskâs Product Access, entitlement, and licensing platforms for cloud and AI products. Shaping enterprise technical strategy, distributed systems, and AI-assisted engineering practices.
đșđž United States â Remote
đ” $196k - $352.1k / year
â° Full Time
đ Senior
đŽ Lead
đ·đ»ââïž Engineer
đ„ 3 hours ago
Strategic Engineer troubleshooting enterprise network and voice issues for Comcast Business, a connectivity and managed-solutions provider. Analyzing OSI-layer faults, configuring call flows, and escalating complex incidents.
đșđž United States â Remote
đ” $26 / hour
â° Full Time
đĄ Mid-level
đ Senior
đ·đ»ââïž Engineer
đŠ H1B Visa Sponsor