Principal Systems Software Engineer – Observability and Telemetry Platform

🔥 27 minutes ago

🏄 California – Remote

info

💵 $272k - $431.3k / year

⏰ Full Time

🔴 Lead

⚙️ Systems Engineer

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Design, implement and support operational and reliability aspects of large scale Observability & Telemetry collection platform with a focus on performance at scale, real time monitoring, logging and alerting • Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation and refinement • Support services before they go live through activities such as system design consulting, developing software tools, platforms and frameworks, capacity management and launch reviews • Maintain services once they are live by measuring and monitoring availability, latency and overall system health • Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity • Practice sustainable incident response and blameless postmortems • Be part of an on call rotation to support production systems

🎯 Requirements

• BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experience • 15+ years of experience with Infrastructure automation, distributed systems design, experience with design, develop tools for running large scale private or public cloud system in Production • 8+ years experience delivering foundational infrastructure and observability platforms • Experience in one or more of the following: Python, Go, Perl or Ruby • In depth knowledge on Linux, Networking and Containers

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🔥 3 hours ago

Netflix

10,000+ employees

📱 Media

👥 B2C

Senior/Staff AI Software Engineer focused on enhancing developer productivity for AIMS at Netflix. Overseeing build, test, and iteration loops to streamline research processes.

🕒 Yesterday

Blue River Technology

201 - 500

🌾 Agriculture

🤖 Artificial Intelligence

🔧 Hardware

Principal Systems Safety Engineer responsible for creating safety frameworks for autonomous systems at Blue River Technology. Leverage engineering insights to develop trust in robotics solutions while collaborating with cross-functional teams.

🕒 2 days ago

Syllo

51 - 200

⚖️ Legal

🤖 Artificial Intelligence

☁️ SaaS

Staff Software Engineer optimizing cloud costs and managing infrastructure for litigation platform at Syllo. Fostering FinOps mindset while contributing to technical design and implementation.

🇺🇸 United States – Remote

💵 $190k - $230k / year

🔥 Funding within the last year

💰 $30M Venture Round - Syllo on 2025-10

⏰ Full Time

🔴 Lead

⚙️ Systems Engineer

🕒 2 days ago

SSM Health

10,000+ employees

🏥 Healthcare

🤝 Non-profit

Enterprise Systems Architect overseeing architecture domains and ensuring alignment with IT strategies. Leading improvements in technology standards and processes for mission-critical systems.

🕒 2 days ago

Carrier

10,000+ employees

🏗️ Construction

🏥 Healthcare

📦 Logistics

Director of Systems Engineering within Carrier focusing on Data Center trends and solutions. Leading product strategy and engineering collaboration for HVAC and control technologies.