AI Storage Solutions Expert

🕒 August 13

🏄 California, Texas – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence

👻 Ghost score 15%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 employees

💼 Consulting

📦 Logistics

🏗️ Construction

💰 Post-IPO Equity on 2023-05

Consulting • Logistics • Construction

Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.

📋 Description

• Deploy and operate parallel/distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre • Design storage architectures optimized for AI checkpoint I/O bursts, sequential dataset reads, and inference KV cache • Implement multi-tenant storage isolation with per-tenant QoS, quotas, and access controls • Configure and optimize GPU Direct Storage for direct GPU-to-storage data paths • Deploy and manage storage networking, including NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX • Diagnose and tune storage performance using IOPS, throughput, latency profiling, fio, IOR, and mdtest • Own runbooks for common storage failure modes • Plan storage capacity for GPU cluster growth and customer workload projections • Manage firmware, data migration, and disaster-recovery procedures • Instrument storage telemetry and feed IO tail latency, checkpoint durations, NVMe SMART, filesystem health, and RDMA counters into the platform metrics/logs/traces store • Partner with the platform team to define the storage-fault predictor, including signals, incident-derived labels, and false-positive tolerances • Convert novel incidents into automation, progressing from SOPs to runbook-as-code and agent-executable remediation • Deliver observability and a baseline predictor for the top three storage-fault classes • Design storage systems for Nvidia GB200-class cluster deployments • Reduce storage-incident MTTR

🎯 Requirements

• 5+ years in enterprise or HPC storage operations • At least 2 years supporting AI/ML workloads • Hands-on deployment and operations experience with at least two of WEKA, VAST Data, Ceph, and DDN/Lustre • Strong understanding of AI training I/O patterns, including checkpoint frequency, dataset loading, and shuffle buffers • Experience with high-performance storage networking, including NFS over RDMA and NVMe-oF • Knowledge of GPU Direct Storage and RDMA-based data transfer • Proficiency in storage performance benchmarking and tuning with fio, IOR, and mdtest • Experience implementing multi-tenant storage with isolation and QoS • Strong Linux systems knowledge, including kernel tuning, filesystem internals, and block device management • Experience shipping an anomaly detector for storage/IO telemetry, or ability to articulate required labels and features • Runbook-as-code mindset; SOPs should be executable by a machine within a quarter

Apply Now

Similar Jobs

🕒 August 13

Agiloft

201 - 500

⚖️ Legal

💼 Consulting

🏥 Healthcare

AI Ops Engineer building governed AI workflows and automation for Agiloft’s data-first contract lifecycle management software. Improving Professional Services implementation efficiency, productivity, and delivery quality.

🇺🇸 United States – Remote

💰 $45M Private Equity Round on 2020-08

⏰ Full Time

🟢 Junior

🟡 Mid-level

🤖 Artificial Intelligence

🦅 H1B Visa Sponsor

infoinfo

🕒 August 12

WelbeHealth

501 - 1000

🏥 Healthcare

⚕️ Healthcare Insurance

AI enablement lead building agents and coaching super users for WelbeHealth, a provider serving vulnerable seniors through PACE. Scaling safe, operationally valuable AI solutions.

🇺🇸 United States – Remote

💵 $120.2k - $158.6k / year

💰 $30M Series C on 2020-02

⏰ Full Time

🟠 Senior

🤖 Artificial Intelligence

🦅 H1B Visa Sponsor

infoinfo

🕒 August 12

Blackbaud

1001 - 5000

☁️ SaaS

🤝 Non-profit

🤝 B2B

Functional Expert Champion helping nonprofits maximize Blackbaud’s RE NXT platform and AI innovations. Driving adoption, customer outcomes, enablement, and product-roadmap influence.

🇺🇸 United States – Remote

💵 $87.7k - $114.2k / year

💰 $6.5M Series A on 2000-02

⏰ Full Time

🟠 Senior

🔴 Lead

🤖 Artificial Intelligence

🕒 August 11

Progressive Leasing

1001 - 5000

💼 Consulting

📦 Logistics

🛒 Retail

AI workforce enablement lead driving adoption and behavior change for Progressive Leasing’s FinTech transformation. Designing manager toolkits, learning programs, and AI-enabled workflow support.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

🤖 Artificial Intelligence

🕒 August 11

Archive

51 - 200

☁️ SaaS

🛍️ eCommerce

🏪 Marketplace

Archive builds branded resale technology for major consumer brands. CX Operations & AI Enablement Specialist automating Zendesk, AI triage, integrations, and reporting.

🇺🇸 United States – Remote

💵 $75k - $100k / year

💰 $30M Series B - Archive on 2025-02

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence