Senior Staff Site Reliability Engineer – Compute Core Engineering

🔥 12 hours ago

🏄 California – Remote

infoinfo

💵 $200k - $322k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Lead initiatives to transform the IT Compute Core Team architecture and build new service offerings across on-premises and cloud environments • Design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP • Build infrastructure for performance and reliability at global scale, including automation, monitoring, high availability, capacity planning, and lifecycle management • Define and implement service-efficiency metrics and drive efficiency through software and hardware optimizations, including SR-IOV and DPU • Apply eBPF and XDP for observability and DDoS mitigation • Collect and review system data for capacity and planning; analyze capacity data and develop enterprise-wide systems plans • Coordinate implementation of infrastructure changes with management personnel • Develop and maintain tools for collecting, analyzing, and visualizing data for reporting, alerting, and monitoring • Collaborate with NVIDIA leadership, senior engineers, program managers, and product managers to develop IT products and services that meet customer needs

🎯 Requirements

• Bachelor’s degree in Engineering, Computer Science, Mathematics, or related field, or equivalent experience • 12+ years of proven experience in compute platform engineering with a focus on automation • Experience designing and deploying containerization architectures and distributed systems infrastructure • Proven experience evaluating application architectures and identifying opportunities for containerization • Strong analytical skills with ability to define and track key performance metrics • Experience developing tools for data analysis and performance profiling • Experience with Terraform and configuration management tools • Proficiency in Go and/or Python • Linux OS proficiency with kernel internals • Experience running large environments consisting of bare-metal build infrastructure • Understanding of network protocols and architectures, including VLAN, VXLAN, SDN, BGP, and Anycast • Deep understanding of infrastructure components such as DNS, LDAP, and security tools • Hands-on experience with containers and their implementation • Experience deploying and managing DNS and LDAP services at scale • Solid understanding of microservices architecture, infrastructure as code (IaC), and configuration management tools

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🔥 13 hours ago

ZoomInfo

1001 - 5000

💼 Consulting

📣 Marketing

☁️ SaaS

Senior DevOps Engineer managing Kubernetes, cloud infrastructure, and data platforms for ZoomInfo’s go-to-market intelligence platform. Improving reliability, automation, observability, and cloud cost efficiency.

🔥 14 hours ago

Bloomerang

201 - 500

💼 Consulting

📣 Marketing

🤝 Non-profit

DevOps Engineer maintaining reliable AWS infrastructure for Bloomerang’s nonprofit giving platform. Optimizing deployments, automation, monitoring, and incident response across multiple technology stacks.

🇺🇸 United States – Remote

💵 $98k - $130k / year

💰 $33M Debt Financing on 2021-02

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 16 hours ago

Summit Racing Equipment

11 - 50

🚘 Automotive

🛒 Retail

🛍️ eCommerce

DevOps Engineer supporting Summit Racing Equipment’s software delivery and reliability systems. Managing monitoring, CI/CD deployments, SDLC processes, and infrastructure-development team integration.

🔥 16 hours ago

PAR Technology

1001 - 5000

🍽️ Food & Beverage

💼 Consulting

📦 Logistics

Senior DevOps Engineer securing cloud infrastructure, Kubernetes, and CI/CD pipelines for PAR Technology’s restaurant technology platform. Driving DevSecOps, observability, compliance, and incident response.

🔥 16 hours ago

Jabil

10,000+ employees

🚘 Automotive

🎖️ Defense

🏥 Healthcare

Lead SRE Security Engineer securing Jabil’s manufacturing, testing, and network infrastructure. Fortinet, Arista, VMware, and cybersecurity leadership for global product solutions.