Senior Storage Production Engineer – DGX Cloud

🔥 12 hours ago

🇦🇺 Australia – Remote

⏰ Full Time

🟠 Senior

🏭 Production Engineer

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Design, implement, and support large-scale storage clusters • Ensure scalability, high availability, and data integrity • Develop and maintain storage monitoring, logging, and alerting systems • Improve storage architectures for AI/ML workloads, including low-latency access, efficient caching, and high-throughput performance • Manage the storage service lifecycle from design and deployment through operation and continuous optimization • Provide system build consulting, automation frameworks, capacity management, and launch reviews • Monitor availability, latency, and system health using predictive analytics and AI-driven automation • Optimize storage through compression, deduplication, tiering, and intelligent workload placement • Scale storage using AI/ML-driven automation, policy-based tiering, and dynamic data migration • Implement encryption, access controls, and auditing mechanisms • Participate in sustainable incident response and blameless root cause analysis • Participate in an on-call rotation supporting storage and production systems

🎯 Requirements

• BS degree or equivalent experience in Computer Science, Storage Systems, or a related technical field • 8+ years of practical experience • Experience with distributed and high-performance storage solutions, including clustered and parallel file systems, distributed object storage, and enterprise-grade storage systems • Understanding of block, file, and object storage technologies, including scalability, reliability, and performance characteristics • Experience with NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics • Expertise in algorithms, data structures, complexity analysis, software design, and automating maintenance of large-scale Linux-based storage systems • Experience with one or more of C/C++, Java, Python, Go, NodeJS, and Bash • Hands-on experience with Ansible, Chef, Puppet, and Terraform • Experience with InfluxDB, Prometheus, Grafana, and the Elastic stack • Excellent written and oral communication skills • Strong work ethic, teamwork, commitment to quality, and task completion • Knowledge of distributed storage systems, replication strategies, erasure coding, capacity planning, performance tuning, and troubleshooting • Experience with Git, code review, pipelines, and CI/CD • Experience with Kubernetes, OpenStack, or hybrid cloud storage architectures • Ability to design automated storage migration, backup, and disaster recovery strategies

🏖️ Benefits

• Full-time employment • Remote work arrangement in Australia

Apply Now

Similar Jobs

🕒 3 days ago

Canva

1001 - 5000

☁️ SaaS

📱 Media

📚 Education

Production Engineering Manager leading Canva’s Embedded team on reliability, resilience and critical production challenges. Shaping engineering practices and developing senior technical leaders.

AWS

Cloud

Google Cloud Platform

Java

Python

C++

Go

🕒 3 days ago

Canva

1001 - 5000

☁️ SaaS

📱 Media

📚 Education

Production Engineering Manager leading Canva’s Core team. Building reliable, scalable engineering capabilities across Canva’s infrastructure and reliability organisation.

AWS

Cloud

Google Cloud Platform

Java

Python

C++

Go

🕒 3 days ago

Canva

1001 - 5000

☁️ SaaS

📱 Media

📚 Education

Engineering Manager leading Canva’s Embedded Production Engineering team. Improving reliability, scalability and resilience across the infrastructure supporting Canva’s 250-million-user design platform.

AWS

Cloud

Google Cloud Platform

Java

Python

C++

Go

🕒 3 days ago

Canva

1001 - 5000

☁️ SaaS

📱 Media

📚 Education

Production Engineering Manager leading Canva’s Core team within its reliability platform. Shaping scalable services, reusable engineering standards and cross-team production capabilities.

AWS

Cloud

Google Cloud Platform

Java

Python

C++

Go

🕒 5 days ago

WP Engine

1001 - 5000

☁️ SaaS

🤝 B2B

Senior Production Engineer leading scalable cloud infrastructure for WP Engine’s WordPress platform. Managing reliability, automation, incidents, and engineering development at internet scale.

Ansible

Apache

AWS

Cloud

Docker

Google Cloud Platform

Kubernetes

Linux

Microservices

MySQL

NGINX

PHP

Prometheus

Python

Terraform

Go