Senior Software Engineer, Resilience Engineering - DGX Cloud

🔥 1 minute ago

🏄 California – Remote

info

💵 $184k - $356.5k / year

⏰ Full Time

🟠 Senior

☁️ Cloud Engineer

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Build org-wide reliability strategy, guiding how NVIDIA matures its operational practices in a 24/7 environment • Stand up a rigorous SLO program, defining and maintaining high standards across teams • Lead incident response for high severity incidents, ensuring low drama and high signal resolution • Build and improve production code daily, enhancing our data platform and related tooling • Implement chaos engineering, failure injection, and resilience testing to elevate our team's standard practices • Improve standards by setting an example with your hands-on experience and leadership

🎯 Requirements

• 8+ years of industry experience • Bachelor's or Master's degree, or equivalent experience operating systems at scale • Strong software engineering skills with current, hands-on experience in Go, Python, or similar languages • Proven experience in establishing and maintaining an SLO program with operational rigor • Practical experience in reliability fields such as chaos engineering and failure injection • Ability to influence across team boundaries through credibility and expertise.

🏖️ Benefits

• Equity • Comprehensive benefits package

Apply Now

Similar Jobs

🔥 2 minutes ago

General Dynamics Information Technology

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Cloud Application Architect responsible for architecting and developing enterprise-class cloud applications. Collaborating with Agile teams to ensure effective implementation and modernization of software solutions.

🔥 14 minutes ago

Trace3

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Cloud Engineer providing engineering and transformation expertise to migrate and modernize environments in cloud infrastructure. Focusing on deploying secure and optimal cloud solutions.

🔥 2 hours ago

DePaul University Athletics

51 - 200

⚽ Sports

📚 Education

Senior Cloud Developer & Analyst focusing on Oracle Cloud ERP, HCM modules management and improvements. Supporting configurations, reports, interfaces, training and partnering with other departments.

🔥 3 hours ago

Accenture Federal Services

10,000+ employees

💼 Consulting

🎖️ Defense

📦 Logistics

Oracle Cloud Conversion developer responsible for converting data in Oracle HCM Cloud. Individual will work with end-to-end business processes and downstream impacts.

🔥 3 hours ago

Chainguard

51 - 200

🔐 Security

☁️ SaaS

🔒 Cybersecurity

Senior Security Engineer securing multi-cloud environments at Chainguard. Architecting cloud security and collaborating across developer and IAM teams to ensure secure operations.