Senior Software Engineer, Infrastructure Automation, Distributed Systems

🕒 July 31

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Design, build, deploy, and run infrastructure services and manage the software life cycle to meet business goals • Participate in defining internal-facing service level objectives and error budgets as part of the observability strategy • Eliminate toil or automate it where the return on investment justifies building and maintaining automation • Practice sustainable blameless incident prevention and incident response • Participate in an on-call rotation • Consult with peer teams on systems design best practices • Provide consultation to peer teams on systems design best practices

🎯 Requirements

• BS degree in Computer Science, a related technical field involving coding, or equivalent experience • 12+ years of relevant experience • Track record of initiating projects, gaining collaboration, and collaborating on others' projects • Experience with infrastructure automation and distributed systems design for large-scale private or public cloud systems in production • Experience with one or more of Python, Go, Perl, or Ruby • In-depth knowledge of one or more of Linux, Networking, Storage, or Containers • Systematic problem-solving approach, strong communication skills, sense of ownership, and drive • Experience using coding assistants, MCP servers, or AI agents to accelerate business impact • Experience with or development of bare metal as a service (BMaaS) systems • Experience with multi-cloud infrastructure services and private or public cloud systems based on Kubernetes, OpenStack, Docker, or Slurm • Experience teaching reliability or cloud systems practices to peers or other companies • Background with NVIDIA Collective Communication Library (NCCL) • No prior experience in a team of any particular name or in an ML/AI-focused team is required

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🕒 July 31

Smithfield Foods

10,000+ employees

🏭 Manufacturing

🌾 Agriculture

🍽️ Food & Beverage

Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.

🕒 July 31

CI&T

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Lead design, deployment, and optimization of NVIDIA cuOpt solutions for AI-focused business operations. Collaborate across teams for effective routing and optimization solutions in Colombia.

🕒 July 31

Rocket.net

11 - 50

☁️ SaaS

🛍️ eCommerce

🏢 Enterprise

Site Reliability Engineer ensuring high standards for servers, services, and customer environments at Rocket.net. Providing advanced technical support and ensuring platform reliability.

🕒 July 31

Branch

501 - 1000

💼 Consulting

📣 Marketing

🔌 API

Senior Site Reliability Engineer at Branch improving platform reliability and scalability with automation. Collaborating with developers to enhance performance and observability of services.

🕒 July 31

Branch

501 - 1000

💼 Consulting

📣 Marketing

🔌 API

Senior Database Reliability Engineer managing MySQL and CloudSQL systems. Ensuring high performance and availability in a remote role at Branch.