Deep Learning Software Engineer, Inference

🔥 14 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🤖 Artificial Intelligence

🎮 Gaming

🚘 Automotive

Artificial Intelligence • Gaming • Automotive

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Performance optimization, analysis, and tuning of DL models in various domains like LLM, Multimodal and Generative AI. • Scale performance of DL models across different architectures and types of NVIDIA accelerators. • Contribute features and code to NVIDIA’s inference libraries, vLLM and SGLang, FlashInfer and LLM software solutions. • Work with cross-collaborative teams across frameworks, NVIDIA libraries and inference optimization innovative solutions.

🎯 Requirements

• Pursuing or recently completed a MS or PhD Computer Engineering, Computer Science, EECS, AI or related field or equivalent experience. • Software development experience. • Excellent C/C++ programming and software design skills. • SW Agile skills are helpful and Python experience is a plus. • Prior experience with training, deploying or optimizing the inference of DL models in production is a plus. • Prior background with performance modeling, profiling, debug, and code optimization or architectural knowledge of CPU and GPU is a plus. • GPU programming experience (CUDA, OAI TRITON or CUTLASS) is a plus.

🏖️ Benefits

• equity • benefits

Apply Now

Similar Jobs

🔥 15 hours ago

Twilio

5001 - 10000

🔌 API

🤝 B2B

Machine Learning Engineer developing scalable, ML-based systems for real-time applications at Twilio. Collaborating with cross-functional teams to drive innovation and deliver personalized customer experiences.

Airflow

AWS

Azure

Cloud

Docker

DynamoDB

Google Cloud Platform

Java

Kafka

Kubernetes

Python

SDLC

Spark

SQL

🕒 Yesterday

Grafana Labs

501 - 1000

🏢 Enterprise

☁️ SaaS

🤖 Artificial Intelligence

Senior Machine Learning Engineer leading the development of personalized recommendation systems for Grafana Labs' Interactive Learning platform. Focus on building applied models and integrating features across teams.

Distributed Systems

GRPC

🕒 3 days ago

Zeitview

51 - 200

⚡ Energy

🤖 Artificial Intelligence

🚗 Transport

Senior MLOps Engineer at Zeitview transforming ML models into reliable production services. Collaborating with scientists, engineers, and DevOps to ensure smooth model deployments and monitoring.

AWS

Cloud

Docker

Kubernetes

Python

Terraform

🕒 3 days ago

Zeitview

51 - 200

⚡ Energy

🤖 Artificial Intelligence

🚗 Transport

Senior MLOps Engineer at Zeitview turning ML models into reliable production-grade services. Collaborating with R&D and DevOps teams on deployment pipelines and cloud infrastructure management.

AWS

Cloud

Docker

GraphQL

Kubernetes

Postgres

Python

Terraform

🕒 3 days ago

Zeitview

51 - 200

⚡ Energy

🤖 Artificial Intelligence

🚗 Transport

Senior MLOps Engineer turning models into reliable services for Zeitview, an intelligent aerial imaging company. Working on infrastructure, pipelines, and tooling for production-grade deployments.

AWS

Cloud

Docker

GraphQL

Kubernetes

Postgres

Python

Terraform