
10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
🕒 June 4
🏄 California, Oregon, +2 more states – Remote
💵 $184k - $356.5k / year
⏰ Full Time
🟠 Senior
🧑💻 Full-stack Engineer
🦅 H1B Visa Sponsor
👻 Ghost score 20%
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
• Lead bring-up, validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how the team operates. • Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks. • Profile and optimize end-to-end workload performance across compute, memory, networking, and communication layers using tools such as Nsight Systems, NCCL tests, and custom microbenchmarks. • Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters, and translate findings into concrete tuning guidance. • Own root-cause analysis of complex failures — hangs, performance regressions, topology sensitivity in large distributed environments. • Define and build the resilience and failure-attribution stack: detecting, triaging, and attributing node, fabric, and workload failures across the cluster at scale. • Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms. • Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams. • Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization. • Mentor engineers, drive technical standards, and act as a force multiplier across the broader performance and infrastructure organization.
• Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience). • 8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including a track record of technical leadership. • Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware. • Deep hands-on experience with NCCL, CUDA-aware distributed execution, and debugging multi-GPU and multi-node workloads at scale. • Proven track record of architecting, debugging, and scaling large-scale distributed systems. • Expert-level Python and C/C++ programming skills. • Experience operating workloads in scheduled, containerized cluster environments. • Excellent analytical, debugging, and communication skills, with the ability to influence across teams.
• equity • benefits
Apply Now🕒 June 4
Senior Software Engineer at Curri developing logistics software. Utilizing AI tooling and end-to-end ownership of projects to enhance last-mile logistics.
🇺🇸 United States – Remote
💵 $185k - $215k / year
💰 Series B - Curri on 2024-07
⏰ Full Time
🟠 Senior
🧑💻 Full-stack Engineer
🕒 June 4
Software Engineer developing innovative tools to support federal clients with cloud-based geospatial data processing. Engaging with scientists and students to enhance earth science data quality.
🇺🇸 United States – Remote
💵 $105k - $141k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
🧑💻 Full-stack Engineer
🕒 June 4
Senior software developer building Azure, .NET, and SQL Server applications for Harvest Valuations’ financial data environment. Remote role limited to Chicagoland residents.
🕒 June 4
Senior Software Engineer owning full-stack product features for Tern’s AI-native travel agency platform. Shipping weekly integrations across Rails, Hotwire, data, mobile, and third-party systems.
🕒 June 4
Senior software engineer architecting scalable AI and cloud-native systems for Trimble’s construction contract-intelligence platform. Leading architecture, DevSecOps, and cross-team engineering across Austin, Texas.
🇺🇸 United States – Remote
💵 $144.6k - $198.8k / year
💰 Post-IPO Debt on 2022-12
⏰ Full Time
🟠 Senior
🧑💻 Full-stack Engineer