Senior Manager, Storage Production Engineering

🔥 36 minutes ago

🏄 California – Remote

info

💵 $272k - $431.3k / year

⏰ Full Time

🟠 Senior

🏭 Production Engineer

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Leading and coaching a team of Storage Production Engineers, creating a collaborative, inclusive, and learning focused environment. • Designing, deploying, and improving large scale storage systems, including distributed storage, parallel file systems, and object storage. • Using automation, monitoring, and analytics to make storage services more reliable, efficient, and easier to operate. • Owning capacity planning, data lifecycle management, cost awareness, and high availability and disaster recovery plans for storage. • Evaluating and adopting modern storage approaches such as NVMe over Fabrics, RDMA, high speed interconnects, and cloud based storage. • Guiding incident response and root cause analysis for storage issues, and putting in place changes that prevent repeat problems. • Partnering with engineering, DevOps, and AI/ML teams to improve data pipelines, access patterns, and workflow performance.

🎯 Requirements

• BS or MS in Computer Science, Storage Systems, or a related technical field, or equivalent experience. • 12+ overall years of experience in large scale storage architecture, operations, production engineering, or infrastructure. • 6+ years of people management or technical leadership experience with storage, infrastructure, or site reliability teams. • Direct experience managing infrastructure operations including on-call rotations, incident response, ongoing maintenance, troubleshooting, and optimization of production systems, managing SLOs and operational KPIs. • Hands on experience with parallel file systems (such as Lustre or GPFS), distributed storage (such as Ceph or MinIO), and enterprise object or NAS platforms (such as S3 compatible systems, NetApp, or Pure Storage). • Strong knowledge of block, file, and object storage, including how to tune performance, protect data, and design for high availability. • Experience with storage networking and protocols like NFS, SMB, iSCSI, Fibre Channel, RDMA, and NVMe-oF. • Practical experience with automation and infrastructure as code using tools such as Terraform, Ansible, or Puppet. • Strong knowledge of monitoring and observability tools (for example Prometheus, InfluxDB, or Elastic stack), logging, and alerting used to operate and improve storage systems.

🏖️ Benefits

• medical, dental, and vision coverage • mental health resources • retirement and 401(k) plans • employee stock purchase plan • paid time off and holidays • family and caregiving leave • range of wellness and development programs

Apply Now

Similar Jobs

🔥 22 hours ago

Marqeta

501 - 1000

💼 Consulting

📦 Logistics

💳 Fintech

Senior Production Support Engineer at Marqeta handling technical support and customer satisfaction. Collaborating with teams to resolve issues and improve service delivery.

🕒 2 days ago

LauraMac

11 - 50

☁️ SaaS

💳 Fintech

🤝 B2B

Software Developer supporting LauraMac’s SaaS platform focused on mortgage efficiency and liquidity. Involves Java and AWS application support in a collaborative team environment.

🕒 2 days ago

Empowers Staffing Inc

11 - 50

💼 Consulting

🎯 Recruiter

🤖 Artificial Intelligence

Machine Learning Engineer designing, deploying, and maintaining scalable ML systems at LAK Technology Inc. Collaborating with teams to operationalize models and implement best practices in MLOps.

🕒 4 days ago

3M Consultancy

1 - 10

💼 Consulting

📣 Marketing

🎯 Recruiter

Production O&M Engineer providing 24x7 application support remotely with focus on incident support and documentation. Requires IRS MBI Clearance and 5+ years of experience in production environments.

🕒 6 days ago

SAIC

10,000+ employees

☁️ SaaS

📣 Marketing

🏢 Enterprise

Systems Engineer providing operational support for complex application cloud systems at SAIC. Detecting and resolving system outages while ensuring compliance with standards and security policies.

🇺🇸 United States – Remote

🔥 Funding within the last year

💰 $500M Post-IPO Debt - SAIC on 2025-09

⏰ Full Time

🟠 Senior

🏭 Production Engineer