Senior Software Engineer, CUDA Deep Learning Systems

đź•’ vor 1 Monat

🏄 California, Texas – Remote

infoinfo

đź’µ $184.000 - $356.500 / Jahr

⏰ Vollzeit

đźź  Senior

🧑‍💻 Full-Stack-Entwickler

🦅 H1B-Visum-Sponsor

infoinfo

đź‘» Geisterscore 1%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of NVIDIA

NVIDIA

10.000+ Mitarbeiter

GegrĂĽndet 1993

🏥 Gesundheitswesen

🏭 Fertigung

🤖 Künstliche Intelligenz

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA ist ein führendes Technologieunternehmen mit Spezialisierung auf beschleunigtes Computing und Künstliche Intelligenz (AI). NVIDIA treibt Fortschritte bei Grafikprozessoren (GPUs), Cloud Computing, Rechenzentren und Virtual Reality voran und fokussiert dabei Branchen wie Gaming, Automotive, Gesundheitswesen und Robotik. Innovationen des Unternehmens wie NVIDIA Omniverse transformieren traditionelle digitale Prozesse, indem sie hochrealistische Simulationen und Rendering-Aufgaben ermöglichen. Die Anwendungen erstrecken sich über zahlreiche Branchen – von autonomen Fahrzeugen mit NVIDIA DRIVE über Gesundheitslösungen mit NVIDIA Clara bis hin zu AI-gestützten Analysen und Workflows.

Beschreibung

• Explore, research, and prototype systems optimizations for advanced deep learning models at the intersection of high-level deep learning frameworks and low-level CUDA through modeling, simulation, and silicon prototyping • Architect and optimize distributed computing systems from single-node to cluster-scale supercomputing environments • Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads • Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines • Collaborate with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design systems and algorithms • Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms • Write clean, effective, and maintainable code and transition prototypes into open-source releases, framework integrations, internal tools, or commercial products

🎯 Anforderungen

• BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience • 8+ years of relevant industry experience or equivalent academic experience after degree achievement • Strong proficiency in C++ and Python programming • Solid background in deep learning fundamentals, with a focus on transformers • Strong understanding of distributed computing, multi-node scaling, and cluster-scale performance challenges • Proven experience in systems programming, computer architecture, and low-level systems performance optimization • Familiarity with GPU accelerator architectures • Hands-on experience with CUDA programming, kernel optimization, and workload profiling • Experience profiling and optimizing generative AI models, including large language models • Research background in machine learning systems or adjacent fields • Experience profiling and optimizing vision models, generative AI architectures, or diffusion models • Track record of initiative and willingness to deep-dive on problems across the stack • Preferred: expertise in performance internals and execution graphs of PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, or Megatron • Preferred: experience with NCCL, MPI, UCX, and distributed machine learning techniques such as pipeline, tensor, or expert parallelism • Preferred: knowledge of numerical methods and low-precision arithmetic such as NVFP4, MXFP4, FP8, or INT8 • Preferred: background in deep learning compilers and ML systems, including Triton, XLA, or torch.compile • Preferred: experience designing agentic AI systems for complex systems and infrastructure problems

🏖️ Vorteile

• Equity • Benefits • Equal opportunity employer • Inclusive work environment

Jetzt Bewerben

Ähnliche Jobs

đź•’ vor 1 Monat

Replicant

51 - 200

đź’Ľ Beratung

🏥 Gesundheitswesen

🛡️ Versicherung

Full-stack engineer building AI voice and chat products for Replicant, whose platform helps enterprise contact centers resolve customer requests with AI. Shipping scalable TypeScript, Node.js, React, and Python systems.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $130.000 - $190.000 / Jahr

⏰ Vollzeit

đźź  Senior

🧑‍💻 Full-Stack-Entwickler

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 1 Monat

Fixify

11 - 50

đź’Ľ Beratung

📦 Logistik

📣 Marketing

Senior Fullstack Engineer building Fixify’s scalable messaging pipelines and user-facing SaaS product. Improving integrations and LLM features through close customer feedback.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

đźź  Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 1 Monat

Netflix

10.000+ Mitarbeiter

📱 Medien

👥 B2C

Full stack engineer building Netflix’s messaging and communications platforms for content production. Architecting scalable frontend and backend systems that power global entertainment workflows.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $388.000 - $558.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

đźź  Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 1 Monat

Alkami Technology

501 - 1000

đź’Ľ Beratung

📣 Marketing

🏦 Bankwesen

Sr Staff Software Engineer building scalable cloud-native banking software for Alkami. Leading architecture, modernization, on-call reliability, and AI-assisted engineering across teams.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $145.000 - $176.000 / Jahr

💰 €300.000.000 Post-IPO Debt - Alkami Technology im 2025-03

⏰ Vollzeit

đźź  Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

đź•’ vor 1 Monat

L3Harris Technologies

10.000+ Mitarbeiter

🏭 Fertigung

đź’Ľ Beratung

📦 Logistik

Full-stack application development manager leading Palantir Foundry enterprise solutions at aerospace and defense technology company L3Harris. Driving workflow automation, data modernization, and operational decision support.

🇺🇸 Vereinigte Staaten – Remote

đź’µ $127.500 - $236.500 / Jahr

⏰ Vollzeit

đźź  Senior

đź”´ Experte

🧑‍💻 Full-Stack-Entwickler

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich