Software Engineer, DGX Cloud AI Infrastructure

🔥 0 minutes ago

🏄 California, Oregon, +2 more states – Remote

infoinfo

💵 $108k - $178.3k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads • Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks • Perform root-cause analysis of failures in large distributed environments • Contribute to resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across the cluster • Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms • Tune runtime settings, communication parameters, and deployment configurations with framework, systems, and platform teams • Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization

🎯 Requirements

• Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience) • Experience developing software for AI, HPC, or systems-level applications • Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution • Background with debugging and scaling distributed systems • Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware • Experience operating workloads in scheduled, containerized cluster environments • Excellent analytical, debugging, and communication skills, and a collaborative approach across teams • Strong Python and C/C++ programming skills • Hands-on experience with NCCL and CUDA-aware distributed execution • Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and InfiniBand / RoCE congestion debugging • Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf • Experience diagnosing performance jitter • Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🔥 5 minutes ago

Garner Health

51 - 200

💼 Consulting

📦 Logistics

🏥 Healthcare

Software Engineer III building AI-powered systems that rank healthcare providers for Garner Health. Remote role using AWS, Kubernetes, and modern software technologies.

🇺🇸 United States – Remote

💵 $175k - $195k / year

⏰ Full Time

🟢 Junior

🟡 Mid-level

🧑‍💻 Full-stack Engineer

🚫👨‍🎓 No degree required

🔥 14 minutes ago

NuScale Power

201 - 500

⚡ Energy

🏭 Manufacturing

Software Developer building control room simulator software for NuScale Power’s nuclear technology. Maintaining production systems, resolving defects, and supporting simulator hardware.

🔥 17 minutes ago

Glydways

51 - 200

🚗 Transport

🤖 Artificial Intelligence

👥 B2C

Senior Embedded Software Engineer developing safety-critical firmware for Glydways’ autonomous transit vehicles. Owning RTOS platforms, hardware bring-up, communication architecture, and automated testing.

🔥 34 minutes ago

Coinbase

1001 - 5000

💼 Consulting

₿ Crypto

💸 Finance

Senior Staff Engineer leading Coinbase’s native Swift/Kotlin mobile platform migration. Architecting AI-assisted iOS and Android development for millions of Coinbase customers.

🔥 34 minutes ago

Coinbase

1001 - 5000

💼 Consulting

₿ Crypto

💸 Finance

Senior mobile engineer rebuilding Coinbase’s retail finance app from React Native to native Swift and Kotlin. Improving performance, architecture, and AI-assisted development for millions of Coinbase customers.