Senior Storage Software Engineer – DGX Cloud

🔥 0 minutes ago

🏄 California – Remote

infoinfo

💵 $224k - $431.3k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Contribute code to open-source parallel and distributed file systems and distributed object storage • Upstream fixes and features and engage with upstream communities and maintainers • Write and review production code as a hands-on storage software lead • Read kernel, NFS, NVMe-oF, or SPDK source to diagnose bugs • Make final technical calls on storage deliveries against measurable targets • Triage, troubleshoot, and root-cause complex storage issues across very large GPU clusters • Investigate I/O and metadata performance, data corruption, and recovery • Validate storage architecture, capabilities, performance, and durability • Run scale tests, benchmarks, and recovery drills • Qualify new builds against measurable performance and durability targets • Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure • Help operators and internal customers apply storage guidelines • Collaborate with training, inference, accelerated-computing, SRE, operations, networking, and security teams • Collaborate with cloud providers, neocloud operators, and storage vendors on common architecture • Use modern AI coding and agentic tools to accelerate building, debugging, validation, and operations

🎯 Requirements

• BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience • Over 12 years of direct experience in storage software engineering • Extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale • Contributions to open-source projects involving a distributed or parallel file system • Hands-on experience writing and reviewing production code, examining file system, kernel, NVMe-oF, or SPDK source, and conducting scale tests or recovery drills • Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance • Strong proficiency in at least one systems language: C, C++, Rust, or Go • Proficiency in Python • Comfortable in Linux kernel storage and networking stacks, including block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, and multipath • Solid understanding of object storage, including S3 / Swift-class • Solid understanding of block storage, including NVMe-oF and iSCSI • Strong written and verbal communication • Comfort operating in a 24/7 production environment • Security-first approach • Maintainers or sustained contributions to widely used public projects • Experience crafting or operating storage for AI training or inference at very large GPU scale • Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience • Kubernetes and CSI driver development for storage • Hands-on experience with SPDK, libfabric, or FUSE performance optimization

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🔥 7 hours ago

Clair

51 - 200

💳 Fintech

🤝 B2B

👥 HR Tech

Senior Engineer building AI-native, scalable products for Clair’s embedded-finance and real-time payroll platform. Owning features end-to-end across React, TypeScript, AWS, and transactional systems.

🕒 Yesterday

ExactCare

501 - 1000

🏥 Healthcare

⚕️ Healthcare Insurance

💊 Pharmaceuticals

Senior Engineer developing and supporting pharmacy-care software at AnewHealth. Building Ruby, JavaScript, Java, and cloud solutions for complex patient needs.

🕒 Yesterday

PBS Radiology Business Experts

51 - 200

💼 Consulting

⚖️ Legal

🛡️ Insurance

C#/.NET Developer modernizing PBS’s Microsoft Access/VBA applications and building enterprise software. Developing maintainable solutions with PostgreSQL in a fully remote IT team.

🕒 Yesterday

ADTRAV Travel Management

201 - 500

💼 Consulting

📦 Logistics

✈️ Travel

Senior Software Engineer building secure, scalable web applications for ADTRAV, a corporate and government travel-management company. Developing full-stack features, cloud services, tests, and production support.

🕒 Yesterday

Zifo

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Technical Lead owning Angular, Python API, and cloud-native product engineering at Zifo. Delivering regulated scientific software for global pharma and biotech customers.