Principal Developer, AI Networking

🕒 June 12

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Characterizing AI workloads and deep learning models aimed at large-scale LLM training and inference on NVIDIA supercomputers. • The role centers on distributed systems with a focus on high-performance networking and NVIDIA communication libraries. • Benchmarking, profiling, and analyzing the performance to find bottlenecks and identify areas for improvement and optimizations, with a strong emphasis on networking aspects. • Developing PyTorch trace-based profiling, analysis, and replaying toolset to aid in benchmarking, debugging, and co-designing network systems for LLM workloads. • Collaborating with multiple teams from hardware to software to provide performance analysis insights. • Defining performance test plans, setting performance expectations for new technologies and solutions, and working to achieve performance targets.

🎯 Requirements

• B.Sc in Computer Science or Software Engineering or equivalent experience. • 15+ years of experience with high-performance networking (RDMA, MPI, NCCL, SHARP). • Demonstrated ability in performance evaluation techniques and approaches. • Experience with NVIDIA GPUs and the CUDA library. • Knowledge of deep learning frameworks like TensorFlow or PyTorch. • Expertise in networking collective communication libraries such as NCCL and protocols like RoCE and RDMA. • Fast and self-learning capabilities with strong analytical and problem-solving skills. • Proficiency in programming languages: Python, Bash, and C++. • Experience with a container-based development environment. • Great teammate who communicates clearly and works well with others.

🏖️ Benefits

• equity • benefits

Apply Now

Similar Jobs

🕒 June 12

Fullsteam

1001 - 5000

🚘 Automotive

🏥 Healthcare

📦 Logistics

Director of Engineering managing high-quality SaaS delivery at Fullsteam, leveraging AI and engineering excellence for growth and modernization. Leading and developing engineering teams with a focus on predictable outcomes.

🕒 June 12

Accenture Federal Services

10,000+ employees

💼 Consulting

🎖️ Defense

📦 Logistics

SAP Fiori Developer leading the design and implementation of user-centric SAP Fiori applications. Collaborating with various teams to ensure business goals and user needs are met.

🕒 June 11

Penn Color, Inc.

501 - 1000

🏭 Manufacturing

🏗️ Construction

Regional Manager overseeing Civil/Site Engineering projects while being fully remote. Driving growth and client relationships for Penn E&R in New Jersey.

🕒 June 10

BlastX Consulting

51 - 200

💼 Consulting

🤝 B2B

🏢 Enterprise

Director of MarTech Engineering at BlastX Consulting, driving digital analytics solutions and client engagement. Leading strategic innovation and team mentorship in a fully remote US-based role.

🕒 June 10

IFS

5001 - 10000

🏗️ Construction

🏭 Manufacturing

📦 Logistics

Principal Product Manager overseeing AI-driven product strategies in Engineering & Construction at IFS. Focused on defining product vision, customer engagement, and delivering measurable business impact.