AI Storage Solutions Expert

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 employees

💼 Consulting

📦 Logistics

🏗️ Construction

💰 Post-IPO Equity on 2023-05

Consulting • Logistics • Construction

Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.

📋 Description

• Deploy and operate parallel/distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre • Design storage architectures optimized for AI checkpoint I/O bursts, sequential dataset reads, and inference KV cache • Implement multi-tenant storage isolation with per-tenant QoS, quotas, and access controls • Configure and optimize GPU Direct Storage for direct GPU-to-storage data paths • Deploy and manage storage networking, including NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX • Diagnose and tune storage performance using IOPS, throughput, latency profiling, fio, IOR, and mdtest • Own runbooks for common storage failure modes • Plan storage capacity for GPU cluster growth and customer workload projections • Manage firmware, data migration, and disaster-recovery procedures • Instrument storage telemetry and feed IO tail latency, checkpoint durations, NVMe SMART, filesystem health, and RDMA counters into the platform metrics/logs/traces store • Partner with the platform team to define the storage-fault predictor, including signals, incident-derived labels, and false-positive tolerances • Convert novel incidents into automation, progressing from SOPs to runbook-as-code and agent-executable remediation • Deliver observability and a baseline predictor for the top three storage-fault classes • Design storage systems for Nvidia GB200-class cluster deployments • Reduce storage-incident MTTR

🎯 Requirements

• 5+ years in enterprise or HPC storage operations • At least 2 years supporting AI/ML workloads • Hands-on deployment and operations experience with at least two of WEKA, VAST Data, Ceph, and DDN/Lustre • Strong understanding of AI training I/O patterns, including checkpoint frequency, dataset loading, and shuffle buffers • Experience with high-performance storage networking, including NFS over RDMA and NVMe-oF • Knowledge of GPU Direct Storage and RDMA-based data transfer • Proficiency in storage performance benchmarking and tuning with fio, IOR, and mdtest • Experience implementing multi-tenant storage with isolation and QoS • Strong Linux systems knowledge, including kernel tuning, filesystem internals, and block device management • Experience shipping an anomaly detector for storage/IO telemetry, or ability to articulate required labels and features • Runbook-as-code mindset; SOPs should be executable by a machine within a quarter

Apply Now

Similar Jobs

🔥 1 hour ago

BD

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

⚕️ Healthcare Insurance

BD Senior Director shaping commercial AI and decision science for a global medical technology company. Leading analytics strategy, governance, predictive modeling, and enterprise AI adoption.

🔥 5 hours ago

Agiloft

201 - 500

⚖️ Legal

💼 Consulting

🏥 Healthcare

AI Ops Engineer building governed AI workflows and automation for Agiloft’s data-first contract lifecycle management software. Improving Professional Services implementation efficiency, productivity, and delivery quality.

🕒 Yesterday

WelbeHealth

501 - 1000

🏥 Healthcare

⚕️ Healthcare Insurance

AI enablement lead building agents and coaching super users for WelbeHealth, a provider serving vulnerable seniors through PACE. Scaling safe, operationally valuable AI solutions.

🕒 Yesterday

Thomson Reuters

10,000+ employees

💼 Consulting

⚖️ Legal

🛡️ Insurance

AI transformation leader turning Thomson Reuters’ trusted legal, tax, compliance, and journalism content and technology into production marketing AI workflows. Leading engineers, governance, and cross-functional delivery.

🕒 Yesterday

Blackbaud

1001 - 5000

☁️ SaaS

🤝 Non-profit

🤝 B2B

Functional Expert Champion helping nonprofits maximize Blackbaud’s RE NXT platform and AI innovations. Driving adoption, customer outcomes, enablement, and product-roadmap influence.

🇺🇸 United States – Remote

💵 $87.7k - $114.2k / year

💰 $6.5M Series A on 2000-02

⏰ Full Time

🟠 Senior

🔴 Lead

🤖 Artificial Intelligence