Senior Site Reliability Engineer, DGX Cloud

🔥 16 hours ago

🏄 California, Texas, +1 more states – Remote

infoinfo

💵 $168k - $333.5k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real-time monitoring, logging, and alerting • Define SLOs/SLIs, monitor error allowances, and streamline reporting • Support services before launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews • Maintain live services by measuring and supervising availability, latency, and overall system health • Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds • Scale systems sustainably through automation and drive changes that improve reliability and velocity • Lead triage and root-cause analysis of high-severity incidents • Practice balanced incident response and blameless postmortems • Participate in on-call rotation to support production services

🎯 Requirements

• BS in Computer Science or related technical field, or equivalent experience • 8+ years of experience operating production services • Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture • Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet) • Proficiency in at least one high-level programming language (e.g., Python, Go) • In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards • Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management • Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc. • Operating GPU-accelerated clusters with KubeVirt in production • Applying generative-AI techniques to reduce operational toil • Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions • Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🔥 16 hours ago

Pluribus Digital

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior DevOps Engineer designing secure, scalable Azure architectures for government agencies. Leading cloud migration, IaC, CI/CD, governance, and federal compliance strategies.

🔥 17 hours ago

FindErnest

11 - 50

💼 Consulting

🏢 Enterprise

🤝 B2B

DevOps Lead managing production Azure AKS environments and advanced Terraform IaC for IT services. Driving Flux CD, Helm, Dynatrace observability, CI/CD, security, and team performance.

🔥 17 hours ago

Guidehouse

10,000+ employees

🏥 Healthcare

🎖️ Defense

📦 Logistics

Cloud infrastructure engineer automating secure AWS and hybrid environments for Guidehouse’s federal clients. Building Terraform, CI/CD, security, monitoring, and disaster recovery capabilities.

🔥 20 hours ago

Qodo (formerly Codium)

11 - 50

🤖 Artificial Intelligence

☁️ SaaS

Senior DevOps Engineer owning Kubernetes-based deployments of Qodo’s AI code review platform in customer AWS, GCP, and Azure environments. Troubleshooting infrastructure and automating enterprise onboarding.

🕒 Yesterday

Custom Mechanical Solutions

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Quality and Reliability Engineer improving field-to-factory quality for Custom Mechanical Solutions’ commercial HVAC equipment. Leading investigations, corrective actions, repair documentation, and reliability improvements across U.S. and Mexican operations.