
10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
🕒 July 31
🏄 California, North Carolina, +3 more states – Remote
💵 $224k - $431.3k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
👻 Ghost score 11%
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
• Design, build, deploy, and run infrastructure services and manage the software life cycle to meet business goals • Participate in defining internal-facing service level objectives and error budgets as part of the observability strategy • Eliminate toil or automate it where the return on investment justifies building and maintaining automation • Practice sustainable blameless incident prevention and incident response • Participate in an on-call rotation • Consult with peer teams on systems design best practices • Provide consultation to peer teams on systems design best practices
• BS degree in Computer Science, a related technical field involving coding, or equivalent experience • 12+ years of relevant experience • Track record of initiating projects, gaining collaboration, and collaborating on others' projects • Experience with infrastructure automation and distributed systems design for large-scale private or public cloud systems in production • Experience with one or more of Python, Go, Perl, or Ruby • In-depth knowledge of one or more of Linux, Networking, Storage, or Containers • Systematic problem-solving approach, strong communication skills, sense of ownership, and drive • Experience using coding assistants, MCP servers, or AI agents to accelerate business impact • Experience with or development of bare metal as a service (BMaaS) systems • Experience with multi-cloud infrastructure services and private or public cloud systems based on Kubernetes, OpenStack, Docker, or Slurm • Experience teaching reliability or cloud systems practices to peers or other companies • Background with NVIDIA Collective Communication Library (NCCL) • No prior experience in a team of any particular name or in an ML/AI-focused team is required
• Equity • Benefits
Apply Now🕒 July 31
Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.
🇺🇸 United States – Remote
💵 $85k - $120k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🕒 July 31
Lead design, deployment, and optimization of NVIDIA cuOpt solutions for AI-focused business operations. Collaborate across teams for effective routing and optimization solutions in Colombia.
🇺🇸 United States – Remote
💰 $5.5M Venture Round on 2014-04
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 July 31
Site Reliability Engineer ensuring high standards for servers, services, and customer environments at Rocket.net. Providing advanced technical support and ensuring platform reliability.
🕒 July 31
DevOps Lead at Neural Earth responsible for maintaining CI/CD pipelines and monitoring AWS cloud infrastructure. Support incident response and operational work for smooth engineering processes.
🇺🇸 United States – Remote
💵 $125k - $156k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 July 31
Forward Deployment Engineer, AI optimizing and deploying AI voice agent solutions for PerfectServe's customers. Collaborating with Product and Customer Success to ensure successful go-live and ongoing support.
🇺🇸 United States – Remote
💵 $120k - $140k / year
💰 Private Equity Round on 2018-05
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)