Senior Deep Learning Software Engineer, Inference

🕒 July 14

🗽 New York, Massachusetts, +2 more states – Remote

infoinfo

💵 $152k - $287.5k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 5%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Performance optimization, analysis, and tuning of DL models in various domains like LLM, Multimodal and Generative AI • Scale performance of DL models across different architectures and types of NVIDIA accelerators • Contribute features and code to NVIDIA’s inference libraries, vLLM and SGLang, FlashInfer and LLM software solutions • Work with cross-collaborative teams across frameworks, NVIDIA libraries and inference optimization innovative solutions

🎯 Requirements

• Masters or PhD or equivalent experience in relevant field (Computer Engineering, Computer Science, EECS, AI) • 5+ years of relevant software development experience • excellent C/C++ programming and software design skills • SW Agile skills are helpful • Python experience is a plus • Prior experience with training, deploying or optimizing the inference of DL models in production is a plus • Prior background with performance modeling, profiling, debug, and code optimization or architectural knowledge of CPU and GPU is a plus • GPU programming experience (CUDA, OAI TRITON or CUTLASS) is a plus

🏖️ Benefits

• health insurance • retirement plans • paid time off • flexible work arrangements • professional development • bonuses • stock options • equity

Apply Now

Similar Jobs

🕒 July 14

Rocket Money (formerly Truebill)

51 - 200

💸 Finance

💳 Fintech

👥 B2C

Senior Full Stack Engineer developing TypeScript and React solutions for Rocket Money's Autopilot technology. Collaborating with teams to optimize banking and investment systems.

🕒 July 14

Diaconia

51 - 200

💼 Consulting

🎖️ Defense

🔒 Cybersecurity

Full-Stack Developer to design, develop, and maintain enterprise-level applications in a mission-critical environment. Seeking skilled professionals in Java Spring Boot and container platforms.

🕒 July 14

Kunai

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Senior Software Engineer focusing on Java services for card network integration in a fintech startup. Collaborate on performance testing and application modernization in a major financial institution.

🕒 July 14

Jahnel Group, an Inc. 5000 company

51 - 200

💼 Consulting

🏥 Healthcare

📣 Marketing

AI Delivery Engineer managing end-to-end AI application delivery for Jahnel Group. Collaborating with clients and engineering teams in a fast-paced environment.

🕒 July 14

Zocdoc

501 - 1000

🏥 Healthcare

⚕️ Healthcare Insurance

🏪 Marketplace

Senior Staff Engineer focusing on vulnerability management at Zocdoc. Collaborating across teams to enhance security and automate remediation processes.