
10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
🔥 0 minutes ago
Ansible
AWS
Azure
Chef
Cloud
Google Cloud Platform
Grafana
Kubernetes
Linux
Microservices
Prometheus
Puppet
Python
Splunk
TCP/IP
Terraform
Go
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
• Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters, focusing on performance at scale, real-time monitoring, logging and alerting • Define SLOs/SLIs, monitor error budgets, and streamline reporting • Support services before launch through system creation consulting, software tools, platforms and frameworks, capacity management, and launch reviews • Maintain live services by measuring and monitoring availability, latency, and overall system health • Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds • Scale systems sustainably through automation and drive changes that improve reliability and velocity • Lead triage and root-cause analysis of high-severity incidents • Practice balanced incident response and blameless postmortems • Participate in an on-call rotation to support production services
• BS in Computer Science or related technical field, or equivalent experience • 10+ years of experience operating production services • Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture • Experience with infrastructure automation tools such as Terraform, Ansible, Chef, or Puppet • Proficiency in at least one high-level programming language such as Python or Go • In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards • Proficient knowledge of SRE principles, including SLOs, SLIs, error budgets, and incident handling • Experience building and operating observability stacks for monitoring, logging, and tracing using tools such as OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, or Splunk • Ability to participate in an on-call rotation • Experience operating production services is required; equivalent experience may substitute for the stated BS credential
Apply Now🕒 July 29
Senior Site Reliability Engineer ensuring reliable service operations at Open Systems with a focus on automation and SRE principles.
🇨🇭 Switzerland – Remote
💰 $1M Venture Round - Open Systems on 2000-07
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
AWS
Azure
Cloud
Distributed Systems
DNS
Google Cloud Platform
GRPC
Kubernetes
Linux
Prometheus
SMTP
TCP/IP
Terraform
Go
🕒 June 13
Site Reliability Engineer at Exoscale responsible for designing and maintaining confidential computing systems. Collaborate on development, improve infrastructure, and enhance performance for a European cloud service provider.
🇨🇭 Switzerland – Remote
💰 Venture Round on 2015-02
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Linux
Python
Go