
10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
🔥 0 minutes ago
🏄 California – Remote
💵 $168k - $270.3k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
• Build tools to improve SRE observability • Participate in the Kubernetes migration journey with VMI setup and problem solving • Rapidly debug and triage incidents and user-reported issues • Automate, script, and develop tooling for new and existing scripts to achieve 100% automation of daily tasks • Support services before launch through system design consulting, software platform and framework development, capacity management, and launch reviews • Participate in an on-call rotation supporting production systems • Drive tools and service development to maintain and improve service SLOs • Partner with Service Owners to drive service reliability • Lead production improvements, including change management, post-mortem reviews, workflow processes, and software automation
• MS or BS in Computer Science, Engineering, or a related field, or equivalent experience • 8+ years of Site Reliability Engineering experience with large-scale distributed microservices in production • Strong Kubernetes background, including complex, highly available VMI setups • Experience leading production improvements, change management, post-mortems, workflow processes, and software automation • Problem-solving and root-cause analysis strengths • Experience with Datadog, Prometheus, Alertmanager, or similar monitoring systems • Experience managing multi-region cloud deployments on AWS, GCP, or Azure • Experience designing and managing deployment pipelines using GitHub Actions, GitLab CI, or ArgoCD • Production-grade coding proficiency in Go, Python, or robust Bash scripting • Required primary production on-call experience responding to and mitigating high-severity infrastructure alerts and service degradations • Excellent communication, presentation, social, and analytical skills • Experience with automated anomaly detection, log clustering tools, or LLM-assisted debugging platforms is a plus • Comfort using AI daily as an SRE • Prior SRE or Service Engineer experience is a plus
• Competitive salary package • Equity • Benefits
Apply Now🔥 33 minutes ago
Senior DevOps/MLOps Engineer operating SimpliGov’s Azure AI platform for government customers. Building secure Kubernetes infrastructure, compliant inference paths, observability, releases, and cost controls.
🔥 52 minutes ago
Senior Backend DevOps Engineer operating AWS containerized microservices for the VA’s JLV clinical data viewer. Building CI/CD, observability, security, and disaster recovery capabilities.
🔥 1 hour ago
Senior Network Reliability Engineer automating insurance company network platforms at Group 1001. Applying SRE, cloud, Kubernetes, security, and observability practices to improve reliability and reduce operational toil.
🔥 5 hours ago
Senior/Staff SRE designing and operating secure on-premises and hybrid-cloud data centers for PathAI’s AI-powered pathology platform. Improving reliability, automation, observability, and incident response for machine-learning infrastructure.
🇺🇸 United States – Remote
💵 $165.8k - $224.4k / year
💰 $165M Series C on 2021-05
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🔥 5 hours ago
DevSecOps Engineer securing cloud products and services for Ardent’s federal national security and defense missions. Automating deployments, vulnerability mitigation, and enterprise system architecture.