
10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
🔥 16 hours ago
🏄 California, Texas, +1 more states – Remote
💵 $168k - $333.5k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
👻 Ghost score 1%
Airflow
Ansible
AWS
Azure
Chef
Cloud
Google Cloud Platform
Grafana
Kubernetes
Linux
Microservices
Prometheus
Puppet
Python
PyTorch
Splunk
TCP/IP
Terraform
Go
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
• Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real-time monitoring, logging, and alerting • Define SLOs/SLIs, monitor error allowances, and streamline reporting • Support services before launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews • Maintain live services by measuring and supervising availability, latency, and overall system health • Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds • Scale systems sustainably through automation and drive changes that improve reliability and velocity • Lead triage and root-cause analysis of high-severity incidents • Practice balanced incident response and blameless postmortems • Participate in on-call rotation to support production services
• BS in Computer Science or related technical field, or equivalent experience • 8+ years of experience operating production services • Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture • Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet) • Proficiency in at least one high-level programming language (e.g., Python, Go) • In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards • Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management • Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc. • Operating GPU-accelerated clusters with KubeVirt in production • Applying generative-AI techniques to reduce operational toil • Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions • Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis
• Equity • Benefits
Apply Now🔥 16 hours ago
Senior DevOps Engineer designing secure, scalable Azure architectures for government agencies. Leading cloud migration, IaC, CI/CD, governance, and federal compliance strategies.
🇺🇸 United States – Remote
💵 $150k - $160k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 17 hours ago
DevOps Lead managing production Azure AKS environments and advanced Terraform IaC for IT services. Driving Flux CD, Helm, Dynatrace observability, CI/CD, security, and team performance.
🇺🇸 United States – Remote
💰 $20k Pre Seed Round - FindErnest on 2023-01
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 17 hours ago
Cloud infrastructure engineer automating secure AWS and hybrid environments for Guidehouse’s federal clients. Building Terraform, CI/CD, security, monitoring, and disaster recovery capabilities.
🇺🇸 United States – Remote
💵 $113k - $188k / year
💰 Grant on 2023-02
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🔥 20 hours ago
Senior DevOps Engineer owning Kubernetes-based deployments of Qodo’s AI code review platform in customer AWS, GCP, and Azure environments. Troubleshooting infrastructure and automating enterprise onboarding.
🇺🇸 United States – Remote
💵 $160k - $180k / year
💰 $10.6M Seed Round on 2023-03
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 Yesterday
Quality and Reliability Engineer improving field-to-factory quality for Custom Mechanical Solutions’ commercial HVAC equipment. Leading investigations, corrective actions, repair documentation, and reliability improvements across U.S. and Mexican operations.
🇺🇸 United States – Remote
💵 $95k - $125k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)