Search Remote Jobs

Service Reliability Engineer

🔥 0 minutes ago

🤠 Texas – Remote

info

đź’µ $168k - $333.5k / year

⏰ Full Time

đźź  Senior

đź”´ Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

đź“‹ Description

• Operate within a 24/7 follow-the-sun support model across multiple continents • Manage a 4-day, 10-hour schedule, including either Saturday or Sunday, with flexible early or late shifts • Monitor and manage extensive production GPU and Kubernetes environments • Detect, prevent, and respond to incidents proactively • Analyze logs, metrics, and system behavior to diagnose issues and implement resolutions • Develop predictive automated support routines • Improve automation through auto-healing and automated break-fix solutions • Perform systems administration, network administration, and security monitoring • Coordinate with domain experts and service owners to resolve complex issues • Continuously improve service quality and operational processes based on incident feedback • Coordinate effectively across teams during incident resolution • Deliver customer-focused support throughout client interactions

🎯 Requirements

• Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management • Familiarity with GPU hardware and high-performance computing environments • Proficiency with Grafana, OpenTelemetry, PagerDuty, and JIRA • Experience with AWS, Azure, GCP, or OCI is a plus; strong preference for on-prem expertise • Ability to work effectively with multifunctional teams • 8+ years of experience coordinating large-scale production systems • More than 3 years of experience in high-availability Internet, Cloud, or Data Center settings • BS in Computer Science, Engineering, Physics, Mathematics, or equivalent experience • Expert-level Linux system administration • Automation using Ansible and/or Python • Strong expertise in shell scripting, DNS, DHCP, storage systems, and core networking • Proven experience maintaining large-scale bare-metal infrastructure • Excellent partnership, documentation, and mentoring skills • Experience with scripting languages, particularly Python • Experience running virtual machines under community-supported or commercial hypervisors • Knowledge of application containers and container orchestration systems • Basic understanding of Git • Ability to master and maintain complicated environments

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🔥 17 minutes ago

Entarian

1001 - 5000

🚀 Aerospace

🎖️ Defense

🏛️ Government

DevSecOps Engineer automating secure cloud infrastructure and CI/CD platforms for Entarian’s U.S. Navy missions. Supporting Kubernetes, Terraform, cloud networking, compliance validation, and infrastructure testing.

🔥 1 hour ago

Endava

10,000+ employees

🏥 Healthcare

📣 Marketing

📦 Logistics

Senior DevOps Engineer designing Dynatrace observability across Endava’s cloud-native enterprise platforms. Improving monitoring, incident response, reliability, and resilience for web, mobile, API, and microservices environments.

🔥 2 hours ago

The Mind Company

51 - 200

🏥 Healthcare

📣 Marketing

đź’Ľ Consulting

Senior DevOps Engineer shaping AI tooling, infrastructure, CI/CD, and mobile releases for a company delivering mental fitness apps. Driving platform reliability and developer productivity across fully remote engineering teams in the Americas.

🔥 5 hours ago

CyberSheath

51 - 200

đź’Ľ Consulting

🎖️ Defense

đź”’ Cybersecurity

Cloud Operations Engineer delivering managed cybersecurity infrastructure services for Defense Industrial Base clients. Supporting Azure, AWS, Linux, Office 365, migrations, and secure technology implementations.

🔥 5 hours ago

JFrog

1001 - 5000

đź’Ľ Consulting

🏥 Healthcare

📦 Logistics

Professional Services DevOps Engineer building CI/CD platforms for JFrog, which manages and secures software delivery from code to production. Designing pipelines, cloud infrastructure and DevOps solutions for enterprise customers.