Systems Software Engineer, Kubernetes Scale

Job not on LinkedIn

🕒 June 26

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Drive end-to-end performance and scale characterization for the NVIDIA DGX Cloud software stack • Collaborate with AI researchers, developers and customers to develop innovative, automated tests • Deep dive into performance and scale issues in complex distributed systems • Design and develop monitoring, reporting and analysis tools for performance and scale testing • Triage, debug and root cause issues related to operating Kubernetes clusters at ultra-large scale • Build and maintain a high-velocity framework that enables continuous performance and scale testing • Document research, methodologies and results clearly and concisely • Engage efficiently with upstream communities

🎯 Requirements

• 2+ years of experience • Computer Architecture, Networking, Storage systems, Accelerators • Bachelors/Masters in Engineering (preferably, Electrical Engineering, Computer Engineering, or Computer Science) or equivalent experience • Expertise in Kubernetes and familiarity with related CNCF projects • Background in working with large scale parallel and distributed accelerator-based systems • Expertise optimizing performance and AI workloads on large scale systems • Experience with performance modeling and benchmarking at scale • Proficiency in Golang/Python • Background with the NVIDIA software ecosystem in both training and inference domains • Expertise with at least one of public CSP infrastructure (GCP, AWS, Azure, OCI for example)

Apply Now

Similar Jobs

🕒 June 24

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Software Engineer developing custom internal security tools to detect, track, and remediate security risks at Akamai. Collaborating with a diverse team to enhance security operations.

Linux

NoSQL

Python

SQL

🕒 June 17

SKELAR

1001 - 5000

📚 Education

🏪 Marketplace

👥 B2C

Backend Engineer specializing in Golang for SKELAR, building a new payment service version from scratch. Collaborating with a dynamic team on innovative tech projects for global markets.

🗣️🇺🇦 Ukrainian Required

Go

🕒 June 17

Madiff

51 - 200

💼 Consulting

📦 Logistics

🏥 Healthcare

.NET Developer for enterprise engineering team delivering modern applications for corporate clients. Focused on backend development using .NET technologies and AI-assisted tools.

SQL

.NET

🕒 June 11

Action1

51 - 200

☁️ SaaS

🔒 Cybersecurity

🏢 Enterprise

Backend Developer designing and coding APIs as part of a development team for high-load applications. Involves backend development, testing, and code optimization at a rapidly growing software company.

AWS

Docker

DynamoDB

EC2

JavaScript

MySQL

Node.js

TypeScript

🕒 June 11

Idego Group

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

Join Idego Group as a programmer. Work with technologies in a pleasant team atmosphere.