HPC Storage Engineer

🕒 Setembro 9

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $180.000 - $260.000 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

🔙 Engenheiro Backend

👻 Score fantasma 5%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Runpod

Runpod

51 - 200 funcionários

Fundada em 2022

🤖 Inteligência Artificial

☁️ SaaS

🤝 B2B

💰 $20.000.000 Seed em 2024-06

Artificial Intelligence • SaaS • B2B

Runpod é uma plataforma de nuvem que fornece computação sob demanda com GPU e infraestrutura gerenciada voltada para desenvolvimento e implantação de IA. A plataforma oferece "Pods" de GPU em 31 regiões globais, endpoints de GPU serverless para inferência de baixa latência, clusters de GPU multinó para treinamento distribuído e um hub para implantar modelos e templates de código aberto. A Runpod destaca-se pelo rápido início (menos de 200ms em iniciações a frio), escalonamento automático de zero a milhares de trabalhadores, suporte para mais de 30 GPU SKUs e ferramentas para todo o ciclo de vida da IA, desde experimentos até a produção, direcionada a desenvolvedores e equipes de IA empresariais.

Descrição

• Own capacity, durability, availability, and performance for network volumes, local NVMe, and S3-compatible object storage • Tune device and filesystem configuration, caching, read-ahead, replication, erasure coding, and client-side mount behavior • Diagnose complex performance problems end to end • Lead capacity expansions, hardware refreshes, migrations, and rebalances without customer-visible disruption • Design and tune high-throughput storage network paths, including MTU, jumbo frames, congestion and flow control, multipath, and NIC/offload configuration • Optimize RDMA/RoCE and high-speed IB/Ethernet fabrics for storage traffic • Collaborate with network engineering on topology, oversubscription, and cross-region data movement • Write production code in Go, Python, or similar for control-plane services, provisioning, data movement, and monitoring • Build and extend control-plane, S3-compatible, CSI, Kubernetes, vendor, and cloud-provider APIs • Automate manual storage operations and manage infrastructure as code • Participate in code review, testing, and CI • Instrument the fleet for IOPS, throughput, latency, errors, retries, utilization, and per-tenant consumption • Build dashboards, SLOs, and alerts • Participate in storage on-call rotations and lead blameless post-incident follow-through • Help determine distributed storage systems, data tiering and placement, network tuning, and petabyte-scale purchasing and deployment

🎯 Requisitos

• 8+ years in infrastructure, storage, or systems engineering, with substantial ownership of production storage at scale • Deep, practical experience with at least one distributed storage system — Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, ZFS-based systems, or comparable • Strong Linux internals and storage-stack knowledge: block layer, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, iSCSI/NVMe-oF • Experience building and/or operating S3-compatible object storage services • Solid networking fundamentals with specific experience tuning networks for storage workloads • Proficiency writing and shipping production code in Go, Python, Rust, or similar • Hands-on experience with observability tooling such as Prometheus, Grafana, or Datadog, including designing metrics • Track record of performance analysis and debugging under real production pressure • Self-starting with general direction • Continuous improvement mindset • Ownership across team boundaries • Collaborative and low-ego, with high confidence • Eligible to work in the United States • Unable to require employment visa sponsorship

🏖️ Benefícios

• Meaningful equity; everyone on the team receives stock options • Generous medical, dental & vision plans • Flexible PTO • Remote-first work environment • Slack-based internal communication • Passionate team on the cutting edge of AI infrastructure • $1,200 Home Office & Equipment Stipend

Candidatar-se

Vagas Similares

🕒 Setembro 9

Miris

11 - 50

☁️ SaaS

🥽 AR/VR

🤝 B2B

Backend Engineer designing scalable Go backend services and APIs for Miris’s global 3D content delivery platform. Improving security, reliability, and performance for AR/VR storytelling.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $102.693 - $287.488 / ano

💰 $26.000.000 Seed Round - MIRIS em 2024-08

⏰ Tempo Integral

🟠 Sênior

🔙 Engenheiro Backend

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 9

GitLab

1001 - 5000

💼 Consultoria

📣 Marketing

🤖 Inteligência Artificial

Staff Backend Engineer building Go-based PostgreSQL automation for GitLab’s DevSecOps orchestration platform. Operating scalable database-as-a-service capabilities across remote teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $152.800 - $259.200 / ano

💰 Secondary Market em 2020-11

⏰ Tempo Integral

🔴 Especialista

🔙 Engenheiro Backend

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 9

Affirm

1001 - 5000

💳 Fintech

👥 B2C

🛍️ Comércio Eletrônico

Backend engineer building scalable APIs that connect partners and merchants to Affirm’s buy-now-pay-later platform. Working with Python or Kotlin, AWS, MySQL, and Kubernetes.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $173.000 - $233.000 / ano

💰 Post-IPO Equity em 2021-01

⏰ Tempo Integral

🟠 Sênior

🔙 Engenheiro Backend

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 8

NewRocket

501 - 1000

💼 Consultoria

🏥 Saúde

🛡️ Seguros

AI Platform Engineer building secure Claude, RAG, and agentic AI infrastructure for NewRocket’s enterprise ServiceNow clients. Operating cloud platforms, LLMOps, integrations, observability, and governance.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

🔙 Engenheiro Backend

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 8

Machinify

1001 - 5000

🏥 Saúde

💼 Consultoria

📦 Logística

Senior backend engineer shaping scalable Java, Scala, and Rust systems for Machinify’s AI-powered healthcare audit products. Guiding agentic engineering and customer-facing application development.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $200.000 - $230.000 / ano

💰 $10.000.000 Series A - Machinify em 2018-10

⏰ Tempo Integral

🟠 Sênior

🔙 Engenheiro Backend

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório