Senior Software Engineer, DGX Cloud AI Infrastructure

🕒 vor 4 Monaten

🏄 California, Oregon, +2 weitere Bundesländer – Remote

infoinfo

💵 $184.000 - $356.500 / Jahr

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🦅 H1B-Visum-Sponsor

infoinfo

👻 Geisterscore 43%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of NVIDIA

NVIDIA

10.000+ Mitarbeiter

Gegründet 1993

🏥 Gesundheitswesen

🏭 Fertigung

🤖 Künstliche Intelligenz

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA ist ein führendes Technologieunternehmen mit Spezialisierung auf beschleunigtes Computing und Künstliche Intelligenz (AI). NVIDIA treibt Fortschritte bei Grafikprozessoren (GPUs), Cloud Computing, Rechenzentren und Virtual Reality voran und fokussiert dabei Branchen wie Gaming, Automotive, Gesundheitswesen und Robotik. Innovationen des Unternehmens wie NVIDIA Omniverse transformieren traditionelle digitale Prozesse, indem sie hochrealistische Simulationen und Rendering-Aufgaben ermöglichen. Die Anwendungen erstrecken sich über zahlreiche Branchen – von autonomen Fahrzeugen mit NVIDIA DRIVE über Gesundheitslösungen mit NVIDIA Clara bis hin zu AI-gestützten Analysen und Workflows.

Beschreibung

• Lead bring-up, validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads • Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks • Profile and optimize end-to-end workload performance across compute, memory, networking, and communication layers using Nsight Systems, NCCL tests, and custom microbenchmarks • Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters • Own root-cause analysis of complex failures, including hangs, performance regressions, and topology sensitivity in large distributed environments • Define and build the resilience and failure-attribution stack for node, fabric, and workload failures across the cluster • Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms • Tune runtime settings, communication parameters, and deployment configurations with framework, systems, and platform teams • Deliver data-driven recommendations based on profiling, benchmark results, and cluster characterization • Mentor engineers, drive technical standards, and influence the broader performance and infrastructure organization

🎯 Anforderungen

• Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience) • 8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including a track record of technical leadership • Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware • Deep hands-on experience with NCCL, CUDA-aware distributed execution, and debugging multi-GPU and multi-node workloads at scale • Proven track record of architecting, debugging, and scaling large-scale distributed systems • Expert-level Python and C/C++ programming skills • Experience operating workloads in scheduled, containerized cluster environments • Excellent analytical, debugging, and communication skills, with the ability to influence across teams • Demonstrated experience debugging and optimizing AI workloads at large scale • Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) • Strong knowledge of GPU cluster fabrics and topology, including NVLink, NVSwitch, PCIe, RoCE, and InfiniBand • Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms • Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure

🏖️ Vorteile

• Equity • Benefits

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 4 Monaten

Curri

51 - 200

📦 Logistik

🏗️ Bauwesen

🚗 Transport

Senior Software Engineer at Curri developing logistics software. Utilizing AI tooling and end-to-end ownership of projects to enhance last-mile logistics.

🇺🇸 Vereinigte Staaten – Remote

💵 $185.000 - $215.000 / Jahr

💰 Series B - Curri im 2024-07

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

Element 84

51 - 200

💼 Beratung

📦 Logistik

🏥 Gesundheitswesen

Software Engineer developing innovative tools to support federal clients with cloud-based geospatial data processing. Engaging with scientists and students to enhance earth science data quality.

🇺🇸 Vereinigte Staaten – Remote

💵 $105.000 - $141.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

Tern

11 - 50

💼 Beratung

📦 Logistik

📣 Marketing

Senior Software Engineer owning full-stack product features for Tern’s AI-native travel agency platform. Shipping weekly integrations across Rails, Hotwire, data, mobile, and third-party systems.

🇺🇸 Vereinigte Staaten – Remote

💵 $175.000 - $200.000 / Jahr

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

LighthouseAI

51 - 200

🏥 Gesundheitswesen

⚖️ Rechtswesen

📦 Logistik

Full Stack Developer building scalable applications for LighthouseAI's pharmaceutical compliance products. Design, develop, and maintain both frontend and backend systems leveraging modern technologies.

🇺🇸 Vereinigte Staaten – Remote

💵 $170.000 - $190.000 / Jahr

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

iCrimeFighter

11 - 50

🏛️ Regierung

🔐 Sicherheit

☁️ SaaS

Senior Full Stack Engineer developing secure APIs and microservices for iCrimeFighter's evidence management platform. Collaborating across stacks with a security-first approach in a fully remote role.

🇺🇸 Vereinigte Staaten – Remote

💰 €125.000 Seed im 2013-12

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich