
11 - 50 employés
Fondée en 2016
🏢 Entreprise
🔒 Cybersecurity
🔧 Matériel
Enterprise • Cybersecurity • Hardware
STN Incorporated est un fournisseur de services gérés en informatique et d'infrastructure cloud de niveau entreprise, offrant une infrastructure sécurisée et prête pour l'audit pour les systèmes critiques pour les entreprises et les charges de travail exigeantes en intelligence artificielle. STN opère un modèle de fonctionnement géré offrant des clouds CPU privés, une infrastructure GPU One AI, un réseau et un stockage sécurisés, ainsi qu'une assistance humaine 24/7 avec des SLA de haute disponibilité. Leurs services incluent l'infrastructure gérée et les opérations cloud, les opérations de cybersécurité et la réponse aux incidents, la gestion de la conformité et des risques (SOC 2 Type II, compatible HIPAA), la sauvegarde et la récupération, et l'acquisition de technologies d'entreprise ainsi que la gestion du cycle de vie. STN sert les entreprises, les entreprises SaaS en forte croissance, les constructeurs d'IA et les développeurs de modèles, les entreprises de robotique/IA physique et les industries réglementées telles que la santé.
🕒 il y a 1 jour
🗣️🇺🇸🇬🇧 Anglais requis
Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

11 - 50 employés
Fondée en 2016
🏢 Entreprise
🔒 Cybersecurity
🔧 Matériel
Enterprise • Cybersecurity • Hardware
STN Incorporated est un fournisseur de services gérés en informatique et d'infrastructure cloud de niveau entreprise, offrant une infrastructure sécurisée et prête pour l'audit pour les systèmes critiques pour les entreprises et les charges de travail exigeantes en intelligence artificielle. STN opère un modèle de fonctionnement géré offrant des clouds CPU privés, une infrastructure GPU One AI, un réseau et un stockage sécurisés, ainsi qu'une assistance humaine 24/7 avec des SLA de haute disponibilité. Leurs services incluent l'infrastructure gérée et les opérations cloud, les opérations de cybersécurité et la réponse aux incidents, la gestion de la conformité et des risques (SOC 2 Type II, compatible HIPAA), la sauvegarde et la récupération, et l'acquisition de technologies d'entreprise ainsi que la gestion du cycle de vie. STN sert les entreprises, les entreprises SaaS en forte croissance, les constructeurs d'IA et les développeurs de modèles, les entreprises de robotique/IA physique et les industries réglementées telles que la santé.
• Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads • Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency • Optimize inference clusters for token generation throughput, low latency, and high GPU utilization • Build and support production AI infrastructure running hundreds to thousands of GPUs • Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers • Perform NCCL benchmarking, analysis, and tuning for collective communication performance • Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS • Configure and tune distributed AI software stacks including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot • Optimize GPU scheduling and resource allocation for training and inference environments • Develop benchmarking and validation processes for hardware, firmware, drivers, and software releases • Identify performance regressions and troubleshoot distributed training issues at scale • Optimize storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O • Work closely with ML engineers to improve training scalability and inference efficiency • Create automation to deploy, validate, benchmark, and monitor GPU clusters • Evaluate emerging AI infrastructure technologies and recommend platform architecture improvements
• 7+ years designing or operating large-scale Linux infrastructure • 5+ years supporting production GPU clusters for AI or HPC workloads • Experience building multi-node GPU training environments from the ground up • Deep expertise with distributed PyTorch training • Extensive experience troubleshooting and optimizing NCCL communications • Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications • Experience benchmarking distributed training with nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf preferred • Understanding of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism • Experience optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth • Experience tuning CUDA, NCCL, UCX, and MPI • Expert-level Linux systems administration skills • Experience with Slurm • Experience using Pyxis and Enroot for containerized GPU workloads • Strong Python and Bash scripting skills • Strong understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet • Experience designing or tuning AI storage, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance • Experience with NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling • Preferred: Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments
Postuler Maintenant🕒 il y a 1 jour
Deployed Engineer partnering with LangChain customers to build, deploy, and operate production AI agents. Designing architectures, leading technical evaluations, and feeding field insights into LangChain’s platform.
🇺🇸 États-Unis – Télétravail
💵 $150 000 - $250 000 / an
💰 €25 000 000 Series A en 2024-02
⏰ Temps Plein
🟡 Intermédiaire
🟠 Senior
👷🏻♀️ Ingénieur
🗣️🇺🇸🇬🇧 Anglais requis
🕒 il y a 1 jour
Principal IAM Engineer Consultant leading secure, scalable CIAM design for regulated financial services clients. Defining authentication architectures, integration patterns, and engineering standards across delivery teams.
🗣️🇺🇸🇬🇧 Anglais requis
🕒 il y a 1 jour
Solutions Engineer configuring AI-powered agents for fund operations at OnCorps, serving asset managers and fund administrators. Leading technical discovery, demonstrations, implementations, and API integrations for customer engagements.
🗣️🇺🇸🇬🇧 Anglais requis
🕒 il y a 1 jour
Fire protection engineer managing designs for hyperscale data centers, federal, mission-critical, and industrial projects at Salas O’Brien. Coordinating technical deliverables, calculations, Revit/BIM modeling, and junior-engineer guidance remotely.
🗣️🇺🇸🇬🇧 Anglais requis
🕒 il y a 1 jour
Distinguished Engineer leading Autodesk’s Product Access, entitlement, and licensing platforms for cloud and AI products. Shaping enterprise technical strategy, distributed systems, and AI-assisted engineering practices.
🗣️🇺🇸🇬🇧 Anglais requis