Search Remote Jobs

Senior Deep Learning Software Infrastructure Engineer

Job not on LinkedIn

🔥 6 hours ago

🏄 California – Remote

info

đź’µ $224k - $431.3k / year

⏰ Full Time

đźź  Senior

đź‘· Infrastructure Engineer

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

đź“‹ Description

• Craft, scale, and harden deep learning infrastructure libraries and frameworks for training on multi-thousand GPU clusters • Improve efficiency throughout the training stack, including data loaders, distributed training, scheduling, and performance monitoring • Build robust training pipelines and libraries for massive video datasets and rapid experimentation • Collaborate with researchers, model engineers, and internal platform teams to enhance efficiency, minimize stalls, and improve training availability • Own core infrastructure components such as orchestration libraries, distributed training frameworks, and fault-resilient training systems • Partner with leadership to scale infrastructure with growing GPU capacity and dataset size while maintaining developer efficiency and stability

🎯 Requirements

• BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, or a related field, or equivalent experience • 12+ years of professional experience building and scaling high-performance distributed systems, ideally in ML, HPC, or large-scale data infrastructure • Extensive knowledge of deep learning frameworks, with PyTorch preferred • Knowledge of large-scale training, including DDP/FSDP, NCCL, tensor parallelism, and pipeline parallelism • Experience with performance profiling • Strong systems background in datacenter networking, including RoCE and IB • Experience with parallel filesystems, including Lustre • Knowledge of storage systems and schedulers such as Slurm and Kubernetes • Proficiency in Python and experience writing production-grade libraries, orchestration layers, and automation tools • Ability to work with ML researchers, infrastructure engineers, and product leads and translate requirements into robust systems • Experience scaling GPU training clusters with more than 1,000 GPUs • Expertise in fault resilience and high availability, including elastic training and large-scale observability • Hands-on technical leadership and ability to establish guidelines for ML systems engineering

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🔥 14 hours ago

SentinelOne

1001 - 5000

đź’Ľ Consulting

🏥 Healthcare

📦 Logistics

Senior AI Platform Engineer owning gateway, Kubernetes, LLM serving, and observability infrastructure. Building AI-native cybersecurity capabilities that protect global enterprises with SentinelOne.

đź•’ Yesterday

Hinshaw & Culbertson LLP

501 - 1000

⚖️ Legal

🛡️ Insurance

🏥 Healthcare

Senior infrastructure engineer designing and supporting servers, virtualization, storage, security, and Azure platforms for national law firm Hinshaw & Culbertson. Leading technical projects, escalations, monitoring, and business continuity initiatives.

đź•’ 2 days ago

HavocAI

11 - 50

📦 Logistics

🏭 Manufacturing

🎖️ Defense

AI Infrastructure Engineer building secure LLM agents, RAG pipelines, and ML tooling. Supporting defense autonomy teams with reliable internal AI systems and workflows.

🇺🇸 United States – Remote

đź’µ $175k - $200k / year

đź’° Seed Round on 2024-09

⏰ Full Time

🟡 Mid-level

đźź  Senior

đź‘· Infrastructure Engineer

đź•’ 2 days ago

Huron

5001 - 10000

🏥 Healthcare

📦 Logistics

📣 Marketing

Senior AI Infrastructure Architect building secure, governed AI platforms for Huron, a global consultancy. Automating cloud infrastructure, model access, telemetry, agent execution, and operational support.

đź•’ 2 days ago

VulnCheck

11 - 50

đź”’ Cybersecurity

🤖 Artificial Intelligence

🏢 Enterprise

Senior Cloud Infrastructure Engineer scaling VulnCheck’s exploit intelligence platform. Building secure cloud infrastructure, CI/CD automation, and SRE practices for cybersecurity products.