Senior Site Reliability Engineer, Production Engineering

🔥 0 minutes ago

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

🏭 Production Engineer

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Lead a global, dynamic, state-of-the-art Service Reliability Operations center • Provide support for NVIDIA Cloud products and services • Partner with Site Reliability Engineering, Security Operations Center, DevOps teams, and other organizations • Support Production Kubernetes Services with a focus on automation and reducing manual tasks • Perform large-scale Kubernetes administration, systems administration, and security monitoring to maintain service SLAs, integrity, and reliability • Use alerts, alarms, and observability tools to monitor, detect, prevent, and respond to incidents • Analyze logs, metrics, and system behavior to troubleshoot issues • Lead root cause analysis and implement effective resolutions • Initiate and lead incident management calls • Coordinate subject matter experts and service owners for timely incident escalation and resolution • Develop monitors, alarms, and alerts to improve service reliability and customer experience

🎯 Requirements

• 7+ years of demonstrated experience administering large-scale production Kubernetes systems in high-availability Internet, Cloud, or Data Center environments • Strong preference for on-prem expertise • BS in Computer Science, Engineering, Mathematics, or equivalent experience • Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management • Familiarity with GPU / DPU hardware and high-performance computing Cluster environments • Strong Linux system administration, DNS, DHCP and core Linux networking (IP Tables, routing, firewalls) experience • Skills to troubleshoot and maintain services on large-scale bare-metal infrastructure • Experience working with CI/CD tools like Jenkins, ArgoCD • Experience in scripting • Programming in Python or Golang or Rust preferred, but not required • Strong communication and soft skills, able to present to cross-functional group members in a persuasive manner

🏖️ Benefits

• 24/7 Production engineering team support • Flexibility to work on split-weekend shifts

Apply Now

Similar Jobs

🕒 June 4

Miratech

501 - 1000

🤝 B2B

💼 Consulting

☁️ SaaS

Production Support Engineer focusing on incident response and operational support in IT and contact center environments at Miratech. Requires 4+ years of experience with tools like ServiceNow and Splunk.

🇮🇳 India – Remote

💰 Private Equity Round on 2022-04

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏭 Production Engineer

Cloud

ServiceNow

Splunk