
10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
🕒 July 31
🏄 California, North Carolina, +3 more states – Remote
💵 $224k - $431.3k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
• Design, build, deploy, and run infrastructure services and manage the software life cycle to meet business goals • Participate in defining internal-facing service level objectives and error budgets as part of the observability strategy • Eliminate toil or automate it where the return on investment justifies building and maintaining automation • Practice sustainable blameless incident prevention and incident response • Participate in an on-call rotation • Consult with peer teams on systems design best practices • Provide consultation to peer teams on systems design best practices
• BS degree in Computer Science, a related technical field involving coding, or equivalent experience • 12+ years of relevant experience • Track record of initiating projects, gaining collaboration, and collaborating on others' projects • Experience with infrastructure automation and distributed systems design for large-scale private or public cloud systems in production • Experience with one or more of Python, Go, Perl, or Ruby • In-depth knowledge of one or more of Linux, Networking, Storage, or Containers • Systematic problem-solving approach, strong communication skills, sense of ownership, and drive • Experience using coding assistants, MCP servers, or AI agents to accelerate business impact • Experience with or development of bare metal as a service (BMaaS) systems • Experience with multi-cloud infrastructure services and private or public cloud systems based on Kubernetes, OpenStack, Docker, or Slurm • Experience teaching reliability or cloud systems practices to peers or other companies • Background with NVIDIA Collective Communication Library (NCCL) • No prior experience in a team of any particular name or in an ML/AI-focused team is required
• Equity • Benefits
Apply Now🕒 July 31
Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.
🕒 July 31
Lead design, deployment, and optimization of NVIDIA cuOpt solutions for AI-focused business operations. Collaborate across teams for effective routing and optimization solutions in Colombia.
🇺🇸 United States – Remote
💰 $5.5M Venture Round on 2014-04
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 July 31
Site Reliability Engineer ensuring high standards for servers, services, and customer environments at Rocket.net. Providing advanced technical support and ensuring platform reliability.
🕒 July 31
Senior Site Reliability Engineer at Branch improving platform reliability and scalability with automation. Collaborating with developers to enhance performance and observability of services.
🇺🇸 United States – Remote
💵 $175k - $185k / year
💰 $282M Series F on 2022-02
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🕒 July 31
Senior Database Reliability Engineer managing MySQL and CloudSQL systems. Ensuring high performance and availability in a remote role at Branch.
🇺🇸 United States – Remote
💵 $175k - $185k / year
💰 $282M Series F on 2022-02
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor