Search Remote Jobs

Senior Manager, Storage Production Engineering

đź•’ July 29

🏄 California – Remote

infoinfo

đź’µ $272k - $431.3k / year

⏰ Full Time

đźź  Senior

🏭 Production Engineer

🦅 H1B Visa Sponsor

infoinfo

đź‘» Ghost score 4%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

đź“‹ Description

• Leading and coaching a team of Storage Production Engineers, creating a collaborative, inclusive, and learning focused environment. • Designing, deploying, and improving large scale storage systems, including distributed storage, parallel file systems, and object storage. • Using automation, monitoring, and analytics to make storage services more reliable, efficient, and easier to operate. • Owning capacity planning, data lifecycle management, cost awareness, and high availability and disaster recovery plans for storage. • Evaluating and adopting modern storage approaches such as NVMe over Fabrics, RDMA, high speed interconnects, and cloud based storage. • Guiding incident response and root cause analysis for storage issues, and putting in place changes that prevent repeat problems. • Partnering with engineering, DevOps, and AI/ML teams to improve data pipelines, access patterns, and workflow performance.

🎯 Requirements

• BS or MS in Computer Science, Storage Systems, or a related technical field, or equivalent experience. • 12+ overall years of experience in large scale storage architecture, operations, production engineering, or infrastructure. • 6+ years of people management or technical leadership experience with storage, infrastructure, or site reliability teams. • Direct experience managing infrastructure operations including on-call rotations, incident response, ongoing maintenance, troubleshooting, and optimization of production systems, managing SLOs and operational KPIs. • Hands on experience with parallel file systems (such as Lustre or GPFS), distributed storage (such as Ceph or MinIO), and enterprise object or NAS platforms (such as S3 compatible systems, NetApp, or Pure Storage). • Strong knowledge of block, file, and object storage, including how to tune performance, protect data, and design for high availability. • Experience with storage networking and protocols like NFS, SMB, iSCSI, Fibre Channel, RDMA, and NVMe-oF. • Practical experience with automation and infrastructure as code using tools such as Terraform, Ansible, or Puppet. • Strong knowledge of monitoring and observability tools (for example Prometheus, InfluxDB, or Elastic stack), logging, and alerting used to operate and improve storage systems.

🏖️ Benefits

• medical, dental, and vision coverage • mental health resources • retirement and 401(k) plans • employee stock purchase plan • paid time off and holidays • family and caregiving leave • range of wellness and development programs

Apply Now

Similar Jobs

đź•’ July 27

Empowers Staffing Inc

11 - 50

đź’Ľ Consulting

🎯 Recruiter

🤖 Artificial Intelligence

Machine Learning Engineer designing, deploying, and maintaining scalable ML systems at LAK Technology Inc. Collaborating with teams to operationalize models and implement best practices in MLOps.

đź•’ June 12

ProSidian Consulting

11 - 50

📦 Logistics

🏭 Manufacturing

🛡️ Insurance

Production Engineer providing technical due diligence and engineering validation for upstream oil and gas projects. Role involves coordinating with various stakeholders and delivering independent engineering advisory services.

đź•’ November 26, 2025

DoubleZero Foundation

1 - 10

đź’Ľ Consulting

📦 Logistics

₿ Crypto

SRE role at DoubleZero focused on automation-first reliability systems in Go, ensuring infrastructure's production readiness and performance.