Incident Response Engineer – Facility Operations Center

🔥 1 hour ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Primary role is to perform coordination and communication across NVIDIA’s datacenter portfolio from an operations perspective regarding incidents, maintenance, and reporting/monitoring. • Develop standards and programs in support of reliability and operations initiatives, including Problem and Change Control, and define and maintain a health score for sites and environments, including testing methods to predict and isolate points of failure, assessing and advising on maintenance strategies, and providing related reporting and metrics. • Study failure data and work with machine learning and AI teams and tools to predict future failures, and facilitate reliability studies such as critical assessments, RAM models, and RCM studies. • Identify and drive automation & process improvement opportunities across catalog quality workflows and reporting. • Coordinate disaster recovery tests, liaise during audits, collaborate with internal partners, and make vital progress to ensure business continuity and compliance. • Perform risk assessments to ensure compliance with policies, procedures, rules & regulations, and data center standards. • Own and present end-to-end key business metrics related to incident response, including ownership and representation of internal and external tooling. • Lead root cause analysis for outages and adjust documentation, workflows, and operating procedures to avoid future incidents. • Assess process improvement & transformation opportunities and partner with process owners & collaborators to scope opportunities, define problem statements and objectives, and structure projects and teams. • Work multi-functionally with other team members and groups within the organization, and develop strong, productive relationships across peer organizations that further the organization's business objectives. • This will incorporate training, coaching, and mentoring Operations teams as needed to empower them to use operations tools and systems to meet daily business needs.

🎯 Requirements

• Bachelor’s degree in a related field (e.g., Electrical Engineering, Mechanical Engineering, Industrial Engineering, Computer Engineering, Telecommunication Engineering, Computer Science, or business-related field) or equivalent experience. • 5+ years of operations or environmental, health, and safety experience within data centers. • Proficient in developing and driving reliability activities (modeling predictions, life cycle testing, stress testing, etc.). • Commercial and financial awareness, with a full comprehension of the impact of failure in translation to business costs, production targets, and fulfillment of customer orders. • Highly developed numeracy, statistical, and reporting skills; ability to analyze, interpret, and apply information, data, and trends. • Enthusiastic about achieving goals and maintaining organization, capable of strategizing and meeting set objectives. • Demonstrated ability to be meticulous, organized, and capable of consolidating data analyses for presentation to large-scale groups. • Proficient in the use of asset database and DCIM solutions to extract data and develop meaningful insights. • Experience in designing, deploying, or maintaining large-scale datacenter infrastructure (whether ACSMEP or networking) or the ability to create strategic infrastructure roadmaps including on-premise, hybrid, and cloud technologies. • Demonstrated knowledge and advanced proficiency working with Microsoft Office Suite software and G-Suite software.

🏖️ Benefits

• Health insurance • Professional development opportunities

Apply Now

Similar Jobs

🕒 Yesterday

DyFlex Solutions

51 - 200

💼 Consulting

📦 Logistics

🏭 Manufacturing

Senior SAP Security Consultant joining DyFlex's Public Cloud practice as a lead security resource for multiple SAP S/4HANA projects. Responsible for client workshops and ongoing security support.

Cloud

Switching

🕒 July 15

DyFlex Solutions

51 - 200

💼 Consulting

📦 Logistics

🏭 Manufacturing

SAP Security Consultant providing hands-on SAP Security support across ECC, S/4HANA, and SAP Public Cloud environments. Collaborating with teams to maintain security and compliance with customer systems.

Cloud

🕒 July 14

CrowdStrike

5001 - 10000

🔒 Cybersecurity

☁️ SaaS

🤖 Artificial Intelligence

Cloud Security Consultant assessing cloud environments for AWS and Azure vulnerabilities. Conducting security assessments, designing detection logic, and collaborating with teams across JAPAC.

AWS

Azure

Cloud

Google Cloud Platform

Python

🕒 July 10

Anthropic

11 - 50

🤖 Artificial Intelligence

☁️ SaaS

🏢 Enterprise

Country Lead overseeing Data Center Security operations in Australia, focusing on security systems and architecture. Collaborate with global teams and manage local security vendors.