HPC Storage Engineer

🔥 12 hours ago

🇺🇸 United States – Remote

💵 $180k - $260k / year

⏰ Full Time

🟠 Senior

🔴 Lead

🔙 Backend Engineer

👻 Ghost score 6%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Runpod

Runpod

51 - 200 employees

Founded 2022

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

💰 $20M Seed on 2024-06

Artificial Intelligence • SaaS • B2B

Runpod is a cloud platform that provides on-demand GPU compute and managed infrastructure tailored for AI development and deployment. It offers GPU "Pods" across 31 global regions, serverless GPU endpoints for low-latency inference, multi-node GPU clusters for distributed training, and a hub for deploying open-source models and templates. Runpod emphasizes fast startup (sub-200ms cold starts), autoscaling from zero to thousands of workers, support for 30+ GPU SKUs, and tooling for the full AI lifecycle from experiment to production, targeting developers and enterprise AI teams.

📋 Description

• Own capacity, durability, availability, and performance for network volumes, local NVMe, and S3-compatible object storage • Tune device and filesystem configuration, caching, read-ahead, replication, erasure coding, and client-side mount behavior • Diagnose complex performance problems end to end • Lead capacity expansions, hardware refreshes, migrations, and rebalances without customer-visible disruption • Design and tune high-throughput storage network paths, including MTU, jumbo frames, congestion and flow control, multipath, and NIC/offload configuration • Optimize RDMA/RoCE and high-speed IB/Ethernet fabrics for storage traffic • Collaborate with network engineering on topology, oversubscription, and cross-region data movement • Write production code in Go, Python, or similar for control-plane services, provisioning, data movement, and monitoring • Build and extend control-plane, S3-compatible, CSI, Kubernetes, vendor, and cloud-provider APIs • Automate manual storage operations and manage infrastructure as code • Participate in code review, testing, and CI • Instrument the fleet for IOPS, throughput, latency, errors, retries, utilization, and per-tenant consumption • Build dashboards, SLOs, and alerts • Participate in storage on-call rotations and lead blameless post-incident follow-through • Help determine distributed storage systems, data tiering and placement, network tuning, and petabyte-scale purchasing and deployment

🎯 Requirements

• 8+ years in infrastructure, storage, or systems engineering, with substantial ownership of production storage at scale • Deep, practical experience with at least one distributed storage system — Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, ZFS-based systems, or comparable • Strong Linux internals and storage-stack knowledge: block layer, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, iSCSI/NVMe-oF • Experience building and/or operating S3-compatible object storage services • Solid networking fundamentals with specific experience tuning networks for storage workloads • Proficiency writing and shipping production code in Go, Python, Rust, or similar • Hands-on experience with observability tooling such as Prometheus, Grafana, or Datadog, including designing metrics • Track record of performance analysis and debugging under real production pressure • Self-starting with general direction • Continuous improvement mindset • Ownership across team boundaries • Collaborative and low-ego, with high confidence • Eligible to work in the United States • Unable to require employment visa sponsorship

🏖️ Benefits

• Meaningful equity; everyone on the team receives stock options • Generous medical, dental & vision plans • Flexible PTO • Remote-first work environment • Slack-based internal communication • Passionate team on the cutting edge of AI infrastructure • $1,200 Home Office & Equipment Stipend

Apply Now

Similar Jobs

🔥 13 hours ago

Miris

11 - 50

☁️ SaaS

🥽 AR/VR

🤝 B2B

Backend Engineer designing scalable Go backend services and APIs for Miris’s global 3D content delivery platform. Improving security, reliability, and performance for AR/VR storytelling.

🇺🇸 United States – Remote

💵 $102.7k - $287.5k / year

💰 $26M Seed Round - MIRIS on 2024-08

⏰ Full Time

🟠 Senior

🔙 Backend Engineer

🔥 15 hours ago

GitLab

1001 - 5000

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

Staff Backend Engineer building Go-based PostgreSQL automation for GitLab’s DevSecOps orchestration platform. Operating scalable database-as-a-service capabilities across remote teams.

🇺🇸 United States – Remote

💵 $152.8k - $259.2k / year

💰 Secondary Market on 2020-11

⏰ Full Time

🔴 Lead

🔙 Backend Engineer

🔥 16 hours ago

General Dynamics Information Technology

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Database Administrator Principal managing AWS PostgreSQL environments, migrations, performance, and data pipelines for GDIT’s government mission customers. Supporting CMS fraud, waste, and abuse investigations.

🔥 18 hours ago

ELOVATE

201 - 500

📦 Logistics

🤝 B2B

🛍️ eCommerce

Senior .NET Developer building full-stack applications and APIs for Elovate’s automated traffic enforcement platform. Applying Microsoft technologies and AI-assisted development to improve safer-road solutions.

🔥 19 hours ago

Affirm

1001 - 5000

💳 Fintech

👥 B2C

🛍️ eCommerce

Backend engineer building scalable APIs that connect partners and merchants to Affirm’s buy-now-pay-later platform. Working with Python or Kotlin, AWS, MySQL, and Kubernetes.