Senior Site Reliability Engineer

🔥 0 minutes ago

🏄 California – Remote

info

💵 $168k - $270.3k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Build tools to improve SRE observability • Participate in the Kubernetes migration journey with VMI setup and problem solving • Rapidly debug and triage incidents and user-reported issues • Automate, script, and develop tooling for new and existing scripts to achieve 100% automation of daily tasks • Support services before launch through system design consulting, software platform and framework development, capacity management, and launch reviews • Participate in an on-call rotation supporting production systems • Drive tools and service development to maintain and improve service SLOs • Partner with Service Owners to drive service reliability • Lead production improvements, including change management, post-mortem reviews, workflow processes, and software automation

🎯 Requirements

• MS or BS in Computer Science, Engineering, or a related field, or equivalent experience • 8+ years of Site Reliability Engineering experience with large-scale distributed microservices in production • Strong Kubernetes background, including complex, highly available VMI setups • Experience leading production improvements, change management, post-mortems, workflow processes, and software automation • Problem-solving and root-cause analysis strengths • Experience with Datadog, Prometheus, Alertmanager, or similar monitoring systems • Experience managing multi-region cloud deployments on AWS, GCP, or Azure • Experience designing and managing deployment pipelines using GitHub Actions, GitLab CI, or ArgoCD • Production-grade coding proficiency in Go, Python, or robust Bash scripting • Required primary production on-call experience responding to and mitigating high-severity infrastructure alerts and service degradations • Excellent communication, presentation, social, and analytical skills • Experience with automated anomaly detection, log clustering tools, or LLM-assisted debugging platforms is a plus • Comfort using AI daily as an SRE • Prior SRE or Service Engineer experience is a plus

🏖️ Benefits

• Competitive salary package • Equity • Benefits

Apply Now

Similar Jobs

🔥 33 minutes ago

SimpliGov

11 - 50

🏛️ Government

☁️ SaaS

⚡ Productivity

Senior DevOps/MLOps Engineer operating SimpliGov’s Azure AI platform for government customers. Building secure Kubernetes infrastructure, compliant inference paths, observability, releases, and cost controls.

🔥 52 minutes ago

VetsEZ

201 - 500

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior Backend DevOps Engineer operating AWS containerized microservices for the VA’s JLV clinical data viewer. Building CI/CD, observability, security, and disaster recovery capabilities.

🔥 1 hour ago

Group 1001

501 - 1000

💼 Consulting

🏥 Healthcare

💸 Finance

Senior Network Reliability Engineer automating insurance company network platforms at Group 1001. Applying SRE, cloud, Kubernetes, security, and observability practices to improve reliability and reduce operational toil.

🔥 5 hours ago

PathAI

501 - 1000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior/Staff SRE designing and operating secure on-premises and hybrid-cloud data centers for PathAI’s AI-powered pathology platform. Improving reliability, automation, observability, and incident response for machine-learning infrastructure.

🔥 5 hours ago

Ardent

51 - 200

💼 Consulting

🎖️ Defense

📦 Logistics

DevSecOps Engineer securing cloud products and services for Ardent’s federal national security and defense missions. Automating deployments, vulnerability mitigation, and enterprise system architecture.