Deep Learning Software Engineer, Inference

🔥 6 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🤖 Artificial Intelligence

🎮 Gaming

Artificial Intelligence • Gaming • Automotive

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Performance optimization, analysis, and tuning of DL models in various domains like LLM, Multimodal and Generative AI. • Scale performance of DL models across different architectures and types of NVIDIA accelerators. • Contribute features and code to NVIDIA’s inference libraries, vLLM and SGLang, FlashInfer and LLM software solutions. • Work with cross-collaborative teams across frameworks, NVIDIA libraries and inference optimization innovative solutions.

🎯 Requirements

• Pursuing or recently completed a MS or PhD Computer Engineering, Computer Science, EECS, AI or related field or equivalent experience. • Software development experience. • Excellent C/C++ programming and software design skills. • SW Agile skills are helpful and Python experience is a plus. • Prior experience with training, deploying or optimizing the inference of DL models in production is a plus. • Prior background with performance modeling, profiling, debug, and code optimization or architectural knowledge of CPU and GPU is a plus. • GPU programming experience (CUDA, OAI TRITON or CUTLASS) is a plus.

🏖️ Benefits

• equity • benefits

Apply Now

Similar Jobs

🔥 15 minutes ago

Twilio

5001 - 10000

Machine Learning Engineer developing scalable, ML-based systems for real-time applications at Twilio. Collaborating with cross-functional teams to drive innovation and deliver personalized customer experiences.

🕒 Yesterday

Grafana Labs

501 - 1000

🏢 Enterprise

☁️ SaaS

🤖 Artificial Intelligence

Senior Machine Learning Engineer leading the development of personalized recommendation systems for Grafana Labs' Interactive Learning platform. Focus on building applied models and integrating features across teams.

🕒 2 days ago

Zeitview

51 - 200

⚡ Energy

🤖 Artificial Intelligence

🚗 Transport

Senior MLOps Engineer at Zeitview transforming ML models into reliable production services. Collaborating with scientists, engineers, and DevOps to ensure smooth model deployments and monitoring.

🕒 2 days ago

Zeitview

51 - 200

⚡ Energy

🤖 Artificial Intelligence

🚗 Transport

Senior MLOps Engineer at Zeitview turning ML models into reliable production-grade services. Collaborating with R&D and DevOps teams on deployment pipelines and cloud infrastructure management.

🕒 2 days ago

Zeitview

51 - 200

⚡ Energy

🤖 Artificial Intelligence

🚗 Transport

Senior MLOps Engineer turning models into reliable services for Zeitview, an intelligent aerial imaging company. Working on infrastructure, pipelines, and tooling for production-grade deployments.