Software Engineer, DGX Cloud AI Infrastructure

🕒 3 dias atrás

🏄 California, Oregon, +2 estados a mais – Remoto

infoinfo

💵 $108.000 - $178.250 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

👻 Score fantasma 1%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of NVIDIA

NVIDIA

10.000+ funcionários

Fundada em 1993

🏥 Saúde

🏭 Manufatura

🤖 Inteligência Artificial

Healthcare • Manufacturing • Artificial Intelligence

A NVIDIA é uma empresa de tecnologia líder, especializada em computação acelerada e inteligência artificial. A companhia é pioneira em avanços em unidades de processamento gráfico (GPUs), computação em nuvem, data centers e realidade virtual, com foco nos setores de games, automotivo, saúde e robótica. As inovações da empresa, como o NVIDIA Omniverse, transformam processos digitais tradicionais ao viabilizar simulações de alta fidelidade e tarefas de renderização. Suas aplicações abrangem diversos setores, desde veículos autônomos com o NVIDIA DRIVE até soluções de saúde com o NVIDIA Clara, além de análises e fluxos de trabalho impulsionados por IA.

Descrição

• Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads • Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks • Perform root-cause analysis of failures in large distributed environments • Contribute to resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across the cluster • Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms • Tune runtime settings, communication parameters, and deployment configurations with framework, systems, and platform teams • Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization

🎯 Requisitos

• Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience) • Experience developing software for AI, HPC, or systems-level applications • Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution • Background with debugging and scaling distributed systems • Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware • Experience operating workloads in scheduled, containerized cluster environments • Excellent analytical, debugging, and communication skills, and a collaborative approach across teams • Strong Python and C/C++ programming skills • Hands-on experience with NCCL and CUDA-aware distributed execution • Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and InfiniBand / RoCE congestion debugging • Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf • Experience diagnosing performance jitter • Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure

🏖️ Benefícios

• Equity • Benefits

Candidatar-se

Vagas Similares

🕒 3 dias atrás

Garner Health

51 - 200

💼 Consultoria

📦 Logística

🏥 Saúde

Software Engineer III building AI-powered systems that rank healthcare providers for Garner Health. Remote role using AWS, Kubernetes, and modern software technologies.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $175.000 - $195.000 / ano

⏰ Tempo Integral

🟢 Júnior

🟡 Pleno

🧑‍💻 Engenheiro Full-stack

🚫👨‍🎓 Sem graduação necessária

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Glydways

51 - 200

🚗 Transporte

🤖 Inteligência Artificial

👥 B2C

Senior Embedded Software Engineer developing safety-critical firmware for Glydways’ autonomous transit vehicles. Owning RTOS platforms, hardware bring-up, communication architecture, and automated testing.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Coinbase

1001 - 5000

💼 Consultoria

₿ Cripto

💸 Finanças

Senior mobile engineer rebuilding Coinbase’s retail finance app from React Native to native Swift and Kotlin. Improving performance, architecture, and AI-assisted development for millions of Coinbase customers.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $186.065 - $218.900 / ano

💰 $21.400.000 Post-IPO Equity em 2022-11

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Coinbase

1001 - 5000

💼 Consultoria

₿ Cripto

💸 Finanças

Senior Staff Engineer leading Coinbase’s native Swift/Kotlin mobile platform migration. Architecting AI-assisted iOS and Android development for millions of Coinbase customers.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $253.895 - $298.700 / ano

💰 $21.400.000 Post-IPO Equity em 2022-11

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Keeper Security, Inc.

501 - 1000

🔒 Cibersegurança

☁️ SaaS

🏢 Corporativo

Senior engineer leading Python/.NET SDKs, CLI tooling, and automation for Keeper Security’s cybersecurity platform. Guiding architecture, hands-on development, and a small engineering team.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 Private Equity Round - Keeper Security em 2023-05

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório