Senior Software Engineer, Fleet Intelligence Agent Systems

🕒 5 days ago

🏄 California, New York – Remote

infoinfo

💵 $152k - $287.5k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Design and develop Fleet Intelligence agent software running on Linux hosts, bare metal systems, and Kubernetes GPU nodes • Build telemetry and health collection for NVIDIA GPUs, DCGM/NVML, drivers, CUDA runtime, InfiniBand, containers, kernel/OS state, CPU, memory, disk, and networking • Develop inventory, enrollment, node identity, local state, attestation, and backend export workflows • Maintain local API, Prometheus metrics, file export, and OTLP/HTTP export paths • Build and improve Kubernetes DaemonSet deployment, systemd service packaging, .deb/.rpm packaging, and container image workflows • Contribute code, tests, documentation, release artifacts, and community-facing engineering practices to open source Fleet Intelligence agent and collector software • Contribute to out-of-band collection using Redfish/BMC interfaces for inventory, GPU attestation, BMC metrics, and logs • Improve collector concurrency, rate limiting, retry behavior, credential handling, partial-failure handling, and backend submission semantics • Collaborate with backend, infrastructure, SRE, security, and datacenter operations teams to deliver reliable GPU fleet observability

🎯 Requirements

• 5+ years of industry software engineering experience • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience • Strong Go development experience for Linux services, CLIs, agents, or daemons • Rust experience, or strong willingness to work in Rust • Strong Linux systems knowledge, including processes, filesystems, networking, service lifecycle, logs, permissions, and host diagnostics • Experience with telemetry, observability, health monitoring, or fleet-management systems • Experience with Docker, Kubernetes, Helm, and production deployment workflows • Experience contributing to open source projects or working in public repositories with code review, issue tracking, documentation, release notes, and signed commits • Familiarity with secure credential handling, enrollment flows, tokens/JWTs, and service-to-service authentication • Strong debugging skills across hardware-adjacent software, operating systems, containers, and distributed backend integrations • Background with NVIDIA datacenter GPUs, DGX systems, DCGM, NVML, CUDA, GPU drivers, XID/SXID events, or GPU diagnostics • Experience building host agents, node agents, collectors, monitoring daemons, or Kubernetes daemonsets • Background with Redfish, BMCs, firmware inventory, secure boot state, PCIe device inventory, or hardware attestation • Experience with OpenTelemetry, Prometheus, OTLP gateways, or metric/log export pipelines • Experience operating software in AI, HPC, cloud, or large-scale datacenter environments • Track record of meaningful open source contributions in systems software, observability, Kubernetes, Linux, hardware telemetry, Rust, or Go ecosystems

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🕒 5 days ago

CACI International Inc

10,000+ employees

💼 Consulting

🎖️ Defense

Full-stack developer building React and ASP.NET systems for CACI’s U.S. Navy Reserve mission. Delivering cloud applications, databases, CI/CD pipelines and automated tests.

🇺🇸 United States – Remote

💵 $75.2k - $158.1k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🕒 6 days ago

Cyberhaven

51 - 200

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

Senior Software Engineer building Cyberhaven’s secure browser extension for data security. Integrating endpoint agents and backend services across Chrome, Edge, and other major browsers.

🇺🇸 United States – Remote

💰 $33M Series B on 2021-12

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🕒 6 days ago

TreviPay

501 - 1000

📦 Logistics

🏭 Manufacturing

✈️ Travel

Software Engineer II building secure, scalable services for TreviPay's global B2B payments and invoicing platform. Developing Java fintech applications, CI/CD pipelines, and high-availability systems.

🕒 6 days ago

Ibrowse Consultoria e Informática

201 - 500

💼 Consulting

👥 B2C

☁️ SaaS

Analista desenvolvedor remoto criando e mantendo aplicações, APIs e soluções de software para consultoria de informática. Atuando em backend, frontend ou fullstack com CI/CD e testes automatizados.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🗣️🇧🇷🇵🇹 Portuguese Required

🕒 6 days ago

Walmart

10,000+ employees

🛒 Retail

🛍️ eCommerce

👥 B2C

Distinguished Software Engineer leading Walmart’s scalable Enterprise Feature Store and cloud-native AI systems. Driving architecture, DevOps, reliability, automation, and engineering mentorship.

🇺🇸 United States – Remote

💵 $130k - $338k / year

💰 $5G Post-IPO Debt - Walmart on 2023-04

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer