Senior Deep Learning Software Infrastructure Engineer

Stelle nicht auf LinkedIn

đź•’ vor 23 Tagen

🏄 California – Remote

infoinfo

đź’µ $224.000 - $431.250 / Jahr

⏰ Vollzeit

đźź  Senior

đź‘· IT-Infrastrukturingenieur

🦅 H1B-Visum-Sponsor

infoinfo

đź‘» Geisterscore 1%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of NVIDIA

NVIDIA

10.000+ Mitarbeiter

GegrĂĽndet 1993

🏥 Gesundheitswesen

🏭 Fertigung

🤖 Künstliche Intelligenz

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA ist ein führendes Technologieunternehmen mit Spezialisierung auf beschleunigtes Computing und Künstliche Intelligenz (AI). NVIDIA treibt Fortschritte bei Grafikprozessoren (GPUs), Cloud Computing, Rechenzentren und Virtual Reality voran und fokussiert dabei Branchen wie Gaming, Automotive, Gesundheitswesen und Robotik. Innovationen des Unternehmens wie NVIDIA Omniverse transformieren traditionelle digitale Prozesse, indem sie hochrealistische Simulationen und Rendering-Aufgaben ermöglichen. Die Anwendungen erstrecken sich über zahlreiche Branchen – von autonomen Fahrzeugen mit NVIDIA DRIVE über Gesundheitslösungen mit NVIDIA Clara bis hin zu AI-gestützten Analysen und Workflows.

Beschreibung

• Craft, scale, and harden deep learning infrastructure libraries and frameworks for training on multi-thousand GPU clusters • Improve efficiency throughout the training stack, including data loaders, distributed training, scheduling, and performance monitoring • Build robust training pipelines and libraries for massive video datasets and rapid experimentation • Collaborate with researchers, model engineers, and internal platform teams to enhance efficiency, minimize stalls, and improve training availability • Own core infrastructure components such as orchestration libraries, distributed training frameworks, and fault-resilient training systems • Partner with leadership to scale infrastructure with growing GPU capacity and dataset size while maintaining developer efficiency and stability

🎯 Anforderungen

• BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, or a related field, or equivalent experience • 12+ years of professional experience building and scaling high-performance distributed systems, ideally in ML, HPC, or large-scale data infrastructure • Extensive knowledge of deep learning frameworks, with PyTorch preferred • Knowledge of large-scale training, including DDP/FSDP, NCCL, tensor parallelism, and pipeline parallelism • Experience with performance profiling • Strong systems background in datacenter networking, including RoCE and IB • Experience with parallel filesystems, including Lustre • Knowledge of storage systems and schedulers such as Slurm and Kubernetes • Proficiency in Python and experience writing production-grade libraries, orchestration layers, and automation tools • Ability to work with ML researchers, infrastructure engineers, and product leads and translate requirements into robust systems • Experience scaling GPU training clusters with more than 1,000 GPUs • Expertise in fault resilience and high availability, including elastic training and large-scale observability • Hands-on technical leadership and ability to establish guidelines for ML systems engineering

🏖️ Vorteile

• Equity • Benefits

Jetzt Bewerben

Ähnliche Jobs

đź•’ vor 23 Tagen

SentinelOne

1001 - 5000

đź’Ľ Beratung

🏥 Gesundheitswesen

📦 Logistik

Senior AI Platform Engineer owning gateway, Kubernetes, LLM serving, and observability infrastructure. Building AI-native cybersecurity capabilities that protect global enterprises with SentinelOne.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $132.000 - $182.000 / Jahr

⏰ Vollzeit

đźź  Senior

đź‘· IT-Infrastrukturingenieur

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 24 Tagen

Hinshaw & Culbertson LLP

501 - 1000

⚖️ Rechtswesen

🛡️ Versicherung

🏥 Gesundheitswesen

Senior infrastructure engineer designing and supporting servers, virtualization, storage, security, and Azure platforms for national law firm Hinshaw & Culbertson. Leading technical projects, escalations, monitoring, and business continuity initiatives.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $120.000 - $145.000 / Jahr

⏰ Vollzeit

đźź  Senior

đź‘· IT-Infrastrukturingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 25 Tagen

HavocAI

11 - 50

📦 Logistik

🏭 Fertigung

🎖️ Verteidigung

AI Infrastructure Engineer building secure LLM agents, RAG pipelines, and ML tooling. Supporting defense autonomy teams with reliable internal AI systems and workflows.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $175.000 - $200.000 / Jahr

đź’° Seed Round im 2024-09

⏰ Vollzeit

🟡 Mittelstufe

đźź  Senior

đź‘· IT-Infrastrukturingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 25 Tagen

Huron

5001 - 10000

🏥 Gesundheitswesen

📦 Logistik

📣 Marketing

Senior AI Infrastructure Architect building secure, governed AI platforms for Huron, a global consultancy. Automating cloud infrastructure, model access, telemetry, agent execution, and operational support.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $175.000 - $245.000 / Jahr

⏰ Vollzeit

đźź  Senior

đź‘· IT-Infrastrukturingenieur

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 26 Tagen

VulnCheck

11 - 50

đź”’ Cybersecurity

🤖 Künstliche Intelligenz

🏢 Unternehmen

Senior Cloud Infrastructure Engineer scaling VulnCheck’s exploit intelligence platform. Building secure cloud infrastructure, CI/CD automation, and SRE practices for cybersecurity products.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

đźź  Senior

đź‘· IT-Infrastrukturingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich