Senior GPU Cloud Storage Solutions Expert – SRE SME

Vaga não está no LinkedIn

🕒 Setembro 11

🏄 California, Texas – Remoto

infoinfo

💵 $180.000 - $320.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 8%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 funcionários

💼 Consultoria

📦 Logística

🏗️ Construção

💰 Post-IPO Equity em 2023-05

Consulting • Logistics • Construction

A Bitdeer Technologies Group (Nasdaq: BTDR) é uma líder na indústria de blockchain e computação de alto desempenho. É uma das maiores detentoras mundiais de hash rate proprietário e fornecedoras de hash rate. A Bitdeer está comprometida em fornecer soluções de computação abrangentes para seus clientes.

Descrição

• Deploy and operate parallel/distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre • Design storage architectures optimized for AI workload patterns such as checkpoint I/O bursts, sequential dataset reads, and KV cache for inference • Implement multi-tenant storage isolation with per-tenant QoS, quotas, and access controls • Configure and optimize GPU Direct Storage for direct GPU-to-storage data paths • Deploy and manage storage networking including NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX for cluster-wide storage orchestration • Diagnose and tune storage performance using IOPS, throughput, latency profiling, fio, IOR, and mdtest • Own the runbook for common failure modes • Plan storage capacity aligned with GPU cluster growth and customer workload projections • Manage firmware, data migration, and disaster recovery procedures • Instrument storage telemetry including IO tail latency, checkpoint durations, NVMe SMART, filesystem health, and RDMA counters • Feed telemetry into the platform team's metrics, logs, and traces store • Partner with the platform team to define the storage-fault predictor, including signals, labels, and false-positive tolerances • Convert novel incidents into automation, progressing from SOPs to runbook-as-code and agent-executable remediation • Deliver observability and a baseline predictor for the top three storage-fault classes • Reduce storage-incident MTTR • Design storage for Nvidia GB200-class clusters

🎯 Requisitos

• 5+ years in enterprise or HPC storage operations, with at least 2 years supporting AI/ML workloads • Hands-on deployment and operations experience with at least two of: WEKA, VAST Data, Ceph, DDN/Lustre • Strong understanding of AI training I/O patterns: checkpoint frequency, dataset loading, shuffle buffers • Experience with high-performance storage networking (NFS over RDMA, NVMe-oF) • Knowledge of GPU Direct Storage and RDMA-based data transfer • Proficiency in storage performance benchmarking and tuning (fio, IOR, mdtest) • Experience implementing multi-tenant storage with isolation and QoS • Strong Linux systems knowledge (kernel tuning, filesystem internals, block device management) • Experience shipping an anomaly detector for storage/IO telemetry or ability to articulate the labels and features needed • Runbook-as-code mindset, with every SOP executable by a machine within a quarter

Candidatar-se

Vagas Similares

🕒 Setembro 11

Paylocity

5001 - 10000

👥 RH Tech

☁️ SaaS

🤝 B2B

DevSecOps Engineer securing Paylocity’s cloud-based HR and payroll software platform. Developing security tooling, integrating build protections, and guiding vulnerability remediation across web and mobile applications.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $96.000 - $130.000 / ano

💰 $10.000.000 Venture Round - Paylocity em 2008-05

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 11

Ookla

201 - 500

📡 Telecomunicações

🏢 Corporativo

Site Reliability Engineer maintaining Ookla’s global cloud, database, and observability infrastructure. Supporting connectivity intelligence services used by hundreds of millions worldwide.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $90.000 - $100.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 10

Zafran Security

51 - 200

🔐 Segurança

Senior DevOps Engineer owning Zafran’s US cybersecurity production environment. Building AWS infrastructure, CI/CD, Kubernetes operations, and production reliability.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 10

Delinea

1001 - 5000

🔒 Cibersegurança

☁️ SaaS

🏢 Corporativo

Senior Site Reliability Engineer owning reliability, monitoring, and incident response for Delinea’s FedRAMP SaaS identity security platform. Automating Azure and AWS operations.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $130.000 - $160.000 / ano

💰 Private Equity Round em 2021-03

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 10

Avanade

10.000+ funcionários

💼 Consultoria

📦 Logística

📣 Marketing

Avanade manager architecting Azure DevOps, GitHub, and AI-enabled software delivery solutions for enterprise clients. Leading DevOps transformation, Copilot adoption, governance, and cloud engineering modernization.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $139.200 - $195.700 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório