Senior Software Engineer, Fleet Intelligence Agent Systems

🔥 0 minutes ago

🏄 California, New York – Remote

infoinfo

💵 $152k - $287.5k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Design and develop Fleet Intelligence agent software running on Linux hosts, bare metal systems, and Kubernetes GPU nodes • Build telemetry and health collection for NVIDIA GPUs, DCGM/NVML, drivers, CUDA runtime, InfiniBand, containers, kernel/OS state, CPU, memory, disk, and networking • Develop inventory, enrollment, node identity, local state, attestation, and backend export workflows • Maintain local API, Prometheus metrics, file export, and OTLP/HTTP export paths • Build and improve Kubernetes DaemonSet deployment, systemd service packaging, .deb/.rpm packaging, and container image workflows • Contribute code, tests, documentation, release artifacts, and community-facing engineering practices to open source Fleet Intelligence agent and collector software • Contribute to out-of-band collection using Redfish/BMC interfaces for inventory, GPU attestation, BMC metrics, and logs • Improve collector concurrency, rate limiting, retry behavior, credential handling, partial-failure handling, and backend submission semantics • Collaborate with backend, infrastructure, SRE, security, and datacenter operations teams to deliver reliable GPU fleet observability

🎯 Requirements

• 5+ years of industry software engineering experience • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience • Strong Go development experience for Linux services, CLIs, agents, or daemons • Rust experience, or strong willingness to work in Rust • Strong Linux systems knowledge, including processes, filesystems, networking, service lifecycle, logs, permissions, and host diagnostics • Experience with telemetry, observability, health monitoring, or fleet-management systems • Experience with Docker, Kubernetes, Helm, and production deployment workflows • Experience contributing to open source projects or working in public repositories with code review, issue tracking, documentation, release notes, and signed commits • Familiarity with secure credential handling, enrollment flows, tokens/JWTs, and service-to-service authentication • Strong debugging skills across hardware-adjacent software, operating systems, containers, and distributed backend integrations • Background with NVIDIA datacenter GPUs, DGX systems, DCGM, NVML, CUDA, GPU drivers, XID/SXID events, or GPU diagnostics • Experience building host agents, node agents, collectors, monitoring daemons, or Kubernetes daemonsets • Background with Redfish, BMCs, firmware inventory, secure boot state, PCIe device inventory, or hardware attestation • Experience with OpenTelemetry, Prometheus, OTLP gateways, or metric/log export pipelines • Experience operating software in AI, HPC, cloud, or large-scale datacenter environments • Track record of meaningful open source contributions in systems software, observability, Kubernetes, Linux, hardware telemetry, Rust, or Go ecosystems

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🔥 4 minutes ago

CACI International Inc

10,000+ employees

💼 Consulting

🎖️ Defense

Full-stack developer building React and ASP.NET systems for CACI’s U.S. Navy Reserve mission. Delivering cloud applications, databases, CI/CD pipelines and automated tests.

🔥 41 minutes ago

CVS Health

10,000+ employees

🏥 Healthcare

⚕️ Healthcare Insurance

🛒 Retail

Senior software engineer supporting CVS Health’s Rebates Application and healthcare technology systems. Driving application stability, testing, scalability, and technical improvements.

🔥 1 hour ago

Cyberhaven

51 - 200

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

Senior Software Engineer building Cyberhaven’s secure browser extension for data security. Integrating endpoint agents and backend services across Chrome, Edge, and other major browsers.

🔥 1 hour ago

CVS Health

10,000+ employees

🏥 Healthcare

⚕️ Healthcare Insurance

🛒 Retail

Senior software engineer developing Java, Spring, and RESTful APIs for CVS Health’s healthcare technology. Leading complex integrations, maintenance, and technical delivery.

🔥 1 hour ago

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Senior Product Manager shaping Cisco Secure Client, Cisco’s unified endpoint security platform. Defining roadmap priorities across AI, connectivity, identity, manageability, and zero trust.