Staff Storage Platform Engineer – AI Storage

Job not on LinkedIn

🔥 1 minute ago

🇺🇸 United States – Remote

⏰ Full Time

🔴 Lead

🏗️ Platform Engineer

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Submer

Submer

51 - 200 employees

Founded 2015

🔧 Hardware

⚡ Energy

🤝 B2B

Hardware • Energy • B2B

Submer is a provider of connected intelligence for AI infrastructure, specializing in advanced liquid/immersion cooling and modular data-center solutions. They design, build and operate AI-ready environments (from power and land to cloud and edge), manufacture immersion cooling pods (SmartPod EVO/EXO), and offer GPUaaS/AIaaS and modular deployment services. Submer emphasizes energy and water savings, high-density thermal architectures for demanding AI workloads, and sovereign-ready, globally deployable solutions.

📋 Description

• Design, build, and operate the AI storage layer powering large-scale GPU infrastructure • Architect scalable storage platforms across edge and core deployments • Define storage strategies for distributed inference, fine-tuning, and training workloads • Design solutions using hyperconverged storage such as StorPool, local NVMe, and disaggregated systems such as VAST and Weka • Define reference architectures, design principles, reusable patterns, and long-term storage roadmaps • Optimize throughput, latency, data locality, resilience, cost, and operability for GPU-heavy clusters • Benchmark and validate storage performance under realistic AI workloads • Implement and maintain CSI drivers and integrate storage with Kubernetes and orchestration systems • Integrate block, object, and shared file storage into the platform • Design multi-tenant, multi-cluster, and multi-site storage environments • Design storage architecture for distributed inference platforms such as NVIDIA Dynamo and llm-d • Optimize KV-cache persistence, token-generation pipelines, and high-concurrency inference workloads • Design high-performance data paths using GPU Direct Storage, RDMA/RoCE, and NVMe-oF • Lead storage performance investigations, incident response, and root-cause analysis • Improve reliability, durability, observability, recovery behavior, and operational standards • Lead end-to-end storage infrastructure delivery from architecture and validation through production rollout • Drive capacity planning, scaling strategies, lifecycle decisions, and safe production changes • Collaborate with compute, networking, platform, DevOps, operations, deployment teams, vendors, and stakeholders • Act as the primary storage design authority and mentor engineers

🎯 Requirements

• Strong hands-on experience designing and operating distributed storage systems for high-performance compute environments • Proven experience designing storage architectures for large-scale AI inference or training platforms, including dataset distribution, checkpointing, and KV-cache storage patterns • Deep knowledge of the Linux storage and I/O stack • Strong understanding of AI workload data access patterns • Experience optimizing storage for GPU-accelerated workloads • Experience with WEKA Data Platform • Familiarity with Kubernetes storage integrations such as CSI • Experience operating large-scale storage clusters • Deep expertise designing and operating storage platforms optimized for GPU-heavy environments and distributed AI workloads • Knowledge of storage hardware, NVMe devices, storage fabrics, and high-performance data paths • Experience designing storage observability systems • Strong automation skills using Python and/or Bash • Experience applying software engineering practices to storage automation and operational tooling • Proven ability to lead complex technical initiatives across teams • Experience with Kubernetes, distributed file systems, object storage, block storage, StorPool, VAST Data, Weka, GPU Direct Storage, RDMA/RoCE, NVMe-oF, and SPDK

🏖️ Benefits

• Attractive compensation package reflecting your expertise and experience • A great work environment characterized by friendliness, international diversity, flexibility, and a hybrid-friendly approach • Exciting career evolution in a fast-growing scale-up • Equal opportunity employer committed to diversity and inclusion

Apply Now

Similar Jobs

🕒 3 days ago

Mondelēz International

10,000+ employees

💼 Consulting

📣 Marketing

📦 Logistics

Platform Engineer securing SAP S/4HANA roles, access, and compliance for Mondelēz International’s global snacking business. Leading GRC, SoD, IAM integration, vulnerability management, and AI automation.

🕒 4 days ago

Zillow

5001 - 10000

🏠 Real Estate

🛍️ eCommerce

👥 B2C

Principal Product Manager shaping Zillow’s internal engineering platform. Defining golden paths, AI-enabled workflows, and adoption strategy for the U.S. real estate platform.

🕒 4 days ago

RxSense

201 - 500

💼 Consulting

📦 Logistics

🏥 Healthcare

Principal Platform Engineer building Kubernetes, Terraform, and CI/CD foundations for RxSense’s pharmacy-benefit technology platform. Enabling secure, observable AI, data, and healthcare workloads.

🕒 4 days ago

Holman

5001 - 10000

🛡️ Insurance

💼 Consulting

🏭 Manufacturing

Director leading digital platform engineering for Holman’s global automotive services organization. Scaling engineering teams and delivering Azure, .NET, React, and AI-enabled platforms.

🕒 5 days ago

Prefect

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Staff Platform Engineer scaling Prefect Cloud’s agentic orchestration platform. Building reliable infrastructure for automation, data, and AI-agent workloads.