Principal Software Engineer, Distributed Systems Engineer – DGX Cloud

🔥 4 hours ago

🌲 North Carolina – Remote

infoinfo

💵 $272k - $431.3k / year

⏰ Full Time

🔴 Lead

⛓️ Blockchain Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Contribute to the DGX Cloud team responsible for production systems enabling large-scale GPU clusters for AI workloads • Develop custom software for scheduling GPU resources on Kubernetes • Implement monitoring and health-management capabilities for reliability, availability, and scalability of GPU assets • Harness data streams from GPU hardware diagnostics, cluster telemetry, and network telemetry • Collaborate with teams across NVIDIA to ensure production AI clusters run reliably, consistently, and with maximum performance • Evaluate system failures and improve services through a defined incident-management process

🎯 Requirements

• Significant software engineering experience with Kubernetes, including cluster operations, operator development, node health monitoring, and GPU resource scheduling • Direct experience in a software engineering role within a highly technical organization, with demonstrable impact • Software development experience with Kubernetes APIs and frameworks, not just operating a cluster • Strong communication skills and ability to work with multifunctional teams, principals, and architects across organizational boundaries and geographies • 15+ years in a similar role and experience with large-scale production systems • Experience with common software engineering principles, tools, and techniques • BS in Computer Science, Engineering, Physics, Mathematics, or a comparable degree, or equivalent experience • Systems programming language proficiency, including Go or Python • Solid understanding of data structures and algorithms • Technical competency managing and automating large-scale distributed systems independent of cloud providers • Advanced hands-on experience and deep understanding of cluster management systems such as Kubernetes, Slurm, or Bright Cluster Manager • Proven operational excellence maintaining reliable and performant AI infrastructure

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🕒 August 15

Fresh Consulting

201 - 500

💼 Consulting

🤖 Artificial Intelligence

🏥 Healthcare

Staff Software Engineer leading distributed scheduling and orchestration for a digital innovation consultancy's Orbital Compute platform. Scaling autonomous compute fleets with secure, reliable, data-sovereign systems.

🕒 July 27

Stradit

11 - 50

🤖 Artificial Intelligence

💳 Fintech

🤝 B2B

Blockchain Architect responsible for designing and overseeing blockchain solutions. Evaluating protocols, ensuring security, and guiding implementation.

🕒 July 23

LiveKit

11 - 50

🔌 API

🤖 Artificial Intelligence

📡 Telecommunications

Staff Software Engineer specializing in distributed systems at LiveKit, designing core services and implementing resilient architectures to support multimodal AI interfaces.

🇺🇸 United States – Remote

💵 $200k - $300k / year

💰 Venture Round on 2022-09

⏰ Full Time

🔴 Lead

⛓️ Blockchain Engineer

🕒 July 13

Netflix

10,000+ employees

📱 Media

👥 B2C

Distributed Systems Engineer designing and maintaining scalable data infrastructure at Netflix. Collaborating with teams to innovate and enhance data-driven solutions.

🕒 June 30

Netflix

10,000+ employees

📱 Media

👥 B2C

Senior technical leader overseeing ad-serving platform architecture and direction. Leveraging distributed systems expertise to optimize performance for high-throughput environments at Netflix.