Senior AI Storage Infrastructure Engineer

🕒 August 19

🏄 California, Texas – Remote

infoinfo

⏰ Full Time

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 17%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 employees

💼 Consulting

📦 Logistics

🏗️ Construction

💰 Post-IPO Equity on 2023-05

Consulting • Logistics • Construction

Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.

📋 Description

• Design, deploy, and maintain CSI drivers for high-performance parallel file systems such as Weka, Lustre, DAOS, and VAST • Architect and implement GPUDirect Storage integrations for DMA between NVMe drives and GPU memory • Develop and manage local NVMe caching strategies for model weights and datasets during distributed training • Optimize IOPS, throughput, and latency across the containerized storage stack • Collaborate with the GPU Systems & Fabric team to optimize storage for RDMA and high-speed interconnects • Implement automated monitoring and alerting for storage performance and mitigate I/O contention or hardware degradation • Define storage policies, quota management, and Kubernetes multi-tenancy isolation strategies • Mentor junior engineers and lead architectural design reviews

🎯 Requirements

• Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field • 5+ years of experience in distributed storage systems and high-performance file systems • Deep understanding of POSIX compliance and file I/O semantics • Deep expertise in the Kubernetes CSI paradigm, including building or extending volume plugins and storage operators • Strong hands-on experience with block/file I/O at the Linux OS level and kernel-level performance tuning • Familiarity with RDMA, InfiniBand, and RoCE and their interaction with storage subsystems • Proven track record of operating, debugging, and scaling large-scale storage environments in production or HPC settings • Experience with Terraform, Ansible, and CI/CD pipelines • Excellent technical communication skills and ability to influence cross-functional architectural decisions • Experience in high-velocity, high-growth engineering environments strongly preferred

🏖️ Benefits

• Equal employment opportunities and non-discrimination protections in accordance with country, state, and local laws

Apply Now

Similar Jobs

🕒 August 17

NetCraftsmen, now BlueAlly

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Senior Cloud Infrastructure Engineer delivering Microsoft cloud migrations and Azure infrastructure for BlueAlly, an IT services provider. Designing identity, endpoint, security, networking, backup, and modernization solutions for public-sector and enterprise clients.

🕒 August 14

fal

51 - 200

🤖 Artificial Intelligence

🔌 API

☁️ SaaS

Kubernetes infrastructure engineer building bare-metal, GPU, storage, and networking environments. Supporting fal’s generative media platform with reliable, scalable customer compute infrastructure.

🕒 August 14

LatamCent

11 - 50

💼 Consulting

📣 Marketing

🎯 Recruiter

Lead security, infrastructure, reliability, and Web3 operations for a global payments and payroll platform. Own GCP, Cloudflare, CI/CD, disaster recovery, and Ethereum settlement security.

🕒 August 10

Preql AI

11 - 50

🤖 Artificial Intelligence

☁️ SaaS

🏢 Enterprise

Forward Deployed Engineer installing and operating Preql’s self-hosted data infrastructure for regulated enterprises. Owning Kubernetes deployments, security reviews, upgrades, and customer production health.

🕒 August 6

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Deep learning infrastructure engineer scaling distributed GPU training for NVIDIA autonomous vehicles. Building resilient pipelines and libraries for massive datasets and rapid model experimentation.