Senior Storage Software Engineer – DGX Cloud

🔥 17 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Contribute to open-source file systems. • Contribute code to open-source parallel and distributed file systems, and distributed object storage. • Upstream fixes and features, and engage directly with the upstream communities and maintainers. • Serve as a hands-on storage software lead. • Write and review production code yourself, and read kernel, NFS, NVMe-oF, or SPDK source when a bug requires it. • Make the final technical calls on storage deliveries against measurable targets. • Triage and troubleshoot at scale. • Triage, troubleshoot, and root-cause large, complex storage issues across very large GPU clusters (tens of thousands of GPUs) — I/O and metadata performance, data corruption, and recovery. • Validate storage architecture, capabilities, performance, and durability. • Run scale tests, benchmarks, and recovery drills, and qualify new builds against measurable performance and durability targets. • Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure, and help operators and internal customers apply them. • Work with training, inference, and accelerated-computing teams, site-reliability and operations, networking, and security, and collaborate with cloud providers, neocloud operators, and storage vendors on a common architecture. • Use modern AI coding and agentic tools day-to-day to accelerate building, debugging, validation, and operations.

🎯 Requirements

• BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience. • Over 12 years of direct experience in storage software engineering, including extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale. • Contributions to open-source projects involving a distributed or parallel file system. • You are fully engaged in engineering tasks. You write and review production code, examine file system, kernel, NVMe-oF, or SPDK source to identify bugs, and personally conduct scale tests or recovery drills instead of assigning them to others. • Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance. • Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python; comfortable in the Linux kernel storage and networking stacks (block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, multipath). • Working knowledge of object storage (S3 / Swift-class) and block storage (NVMe-oF, iSCSI). • Strong written and verbal communication; capable of clarifying complex technical trade-offs to engineers, SREs, vendors, and internal customers. • Comfort operating in a 24/7 production environment where storage incidents directly impact GPU availability, with a security-first approach baked into every build. • 100% hands-on engineering. You write and review production code, read file system, kernel, NVMe-oF, or SPDK source to chase bugs, and run scale tests or recovery drills yourself rather than delegating.

🏖️ Benefits

• equity • benefits

Apply Now

Similar Jobs

🔥 19 hours ago

bswift

1001 - 5000

🏥 Healthcare

💼 Consulting

🛡️ Insurance

Sr. Database Engineer (Cloud DBA) managing SQL Server environments and cloud operations at bswift. Collaborating with teams to enhance database reliability and support business growth.

🇺🇸 United States – Remote

💵 $140k - $160k / year

💰 $51M Private Equity Round on 2014-04

⏰ Full Time

🟠 Senior

☁️ Cloud Engineer

AWS

Cloud

EC2

SQL

Terraform

🔥 20 hours ago

Shadeform

11 - 50

☁️ SaaS

🏪 Marketplace

🤝 B2B

Senior Software Engineer building and maintaining the Shadeform GPU platform and AI infrastructure services. Working on novel solutions for challenges in the GPU market.

Cloud

Kubernetes

🕒 Yesterday

Southern Bancorp

201 - 500

🏦 Banking

🌍 Social Impact

Cloud Cybersecurity Engineer securing Microsoft 365 and Azure environments through secure architecture and identity controls. Balancing data protection with functionality to ensure operational excellence.

🇺🇸 United States – Remote

💰 $4.9M Grant - Southern Bancorp on 2023-04

⏰ Full Time

🟡 Mid-level

🟠 Senior

☁️ Cloud Engineer

Azure

Cloud

🕒 Yesterday

Vestis Corporation

10,000+ employees

🤝 B2B

🏢 Enterprise

Senior Director overseeing technical direction of cloud and infrastructure platform architecture. Leading teams in developing and executing modern infrastructure strategies in a hybrid environment.

Azure

Cloud

Cyber Security

Oracle

SQL

🕒 Yesterday

Corpay

10,000+ employees

💳 Fintech

🤝 B2B

☁️ SaaS

AWS Cloud Engineer optimizing cloud solutions for Corpay's Domestic Payables division. Driving technical efforts involving design, development, and implementation of AWS applications.

Angular

AWS

Cloud

ETL

JavaScript

Node.js