
201 - 500 employees
💼 Consulting
📦 Logistics
🏗️ Construction
💰 Post-IPO Equity on 2023-05
Consulting • Logistics • Construction
Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.
🔥 50 minutes ago
🏄 California, Texas – Remote
💵 $180k - $320k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
👻 Ghost score 4%
Linux
NFS
Improve your chances of getting an interview by checking your resume score before you apply.

201 - 500 employees
💼 Consulting
📦 Logistics
🏗️ Construction
💰 Post-IPO Equity on 2023-05
Consulting • Logistics • Construction
Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.
• Deploy and operate parallel/distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre • Design storage architectures optimized for AI workload patterns such as checkpoint I/O bursts, sequential dataset reads, and KV cache for inference • Implement multi-tenant storage isolation with per-tenant QoS, quotas, and access controls • Configure and optimize GPU Direct Storage for direct GPU-to-storage data paths • Deploy and manage storage networking including NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX for cluster-wide storage orchestration • Diagnose and tune storage performance using IOPS, throughput, latency profiling, fio, IOR, and mdtest • Own the runbook for common failure modes • Plan storage capacity aligned with GPU cluster growth and customer workload projections • Manage firmware, data migration, and disaster recovery procedures • Instrument storage telemetry including IO tail latency, checkpoint durations, NVMe SMART, filesystem health, and RDMA counters • Feed telemetry into the platform team's metrics, logs, and traces store • Partner with the platform team to define the storage-fault predictor, including signals, labels, and false-positive tolerances • Convert novel incidents into automation, progressing from SOPs to runbook-as-code and agent-executable remediation • Deliver observability and a baseline predictor for the top three storage-fault classes • Reduce storage-incident MTTR • Design storage for Nvidia GB200-class clusters
• 5+ years in enterprise or HPC storage operations, with at least 2 years supporting AI/ML workloads • Hands-on deployment and operations experience with at least two of: WEKA, VAST Data, Ceph, DDN/Lustre • Strong understanding of AI training I/O patterns: checkpoint frequency, dataset loading, shuffle buffers • Experience with high-performance storage networking (NFS over RDMA, NVMe-oF) • Knowledge of GPU Direct Storage and RDMA-based data transfer • Proficiency in storage performance benchmarking and tuning (fio, IOR, mdtest) • Experience implementing multi-tenant storage with isolation and QoS • Strong Linux systems knowledge (kernel tuning, filesystem internals, block device management) • Experience shipping an anomaly detector for storage/IO telemetry or ability to articulate the labels and features needed • Runbook-as-code mindset, with every SOP executable by a machine within a quarter
Apply Now🔥 3 hours ago
DevOps Engineer 4 building secure CI/CD and cloud platforms for Centex Technologies’ programs and customers. Automating infrastructure, observability, security, and reliable software delivery.
🔥 3 hours ago
DevSecOps Engineer securing Paylocity’s cloud-based HR and payroll software platform. Developing security tooling, integrating build protections, and guiding vulnerability remediation across web and mobile applications.
🇺🇸 United States – Remote
💵 $96k - $130k / year
💰 $10M Venture Round - Paylocity on 2008-05
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 6 hours ago
Senior DevOps Engineer owning AWS infrastructure, Kubernetes, CI/CD, and observability for EIS SaaS applications. Building secure, reliable end-to-end delivery platforms.
🇺🇸 United States – Remote
💵 $125k - $140k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 15 hours ago
Site Reliability Engineer maintaining Ookla’s global cloud, database, and observability infrastructure. Supporting connectivity intelligence services used by hundreds of millions worldwide.
🇺🇸 United States – Remote
💵 $90k - $100k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 17 hours ago
Senior DevOps Engineer owning Zafran’s US cybersecurity production environment. Building AWS infrastructure, CI/CD, Kubernetes operations, and production reliability.