
201 - 500 employees
💼 Consulting
📦 Logistics
🏗️ Construction
💰 Post-IPO Equity on 2023-05
Consulting • Logistics • Construction
Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.
🔥 0 minutes ago
🏄 California, Texas – Remote
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Improve your chances of getting an interview by checking your resume score before you apply.

201 - 500 employees
💼 Consulting
📦 Logistics
🏗️ Construction
💰 Post-IPO Equity on 2023-05
Consulting • Logistics • Construction
Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.
• Operate InfiniBand and RoCEv2 fabrics carrying NCCL traffic across Bitdeer's US data centers • Manage fat-tree, rail-optimized, and dragonfly network topologies for GPU clusters of 100–10,000 GPUs • Monitor and tune IB and RoCE performance, including adaptive routing, congestion control, and traffic isolation • Tune NCCL communication through topology detection, ring/tree algorithm selection, and GDR configuration • Manage firmware lifecycle across InfiniBand switches and HCAs • Diagnose link flaps, symbol errors, packet drops, routing anomalies, and credit stalls • Coordinate with Nvidia/Mellanox support on escalations, bugs, and RMAs • Feed IB/RoCE telemetry into the AIOps platform's collection pipeline • Partner with the platform team to define link and straggler predictors • Convert incidents into labeled examples for the fault-prediction engine • Turn routine mitigations into workflows for automated remediation
• 5+ years in data center networking, including at least 3 years focused on InfiniBand or RoCE fabrics • Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale • Strong understanding of IB subnet management, partitioning, and QoS • Experience deploying RoCEv2, including PFC, ECN, and DCQCN configuration • Proficiency with UFM or equivalent IB fabric management tools • Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices • Experience diagnosing IB/RoCE issues using ibdiagnet, perfquery, ibstat, and similar tools • Understanding of NCCL and GPU communication mapping to network topology • Telemetry-driven operations experience, such as building dashboards or alerts on RDMA counters at scale, or ability to define features for a fabric-health model • Runbook-as-code mindset
Apply Now🔥 34 minutes ago
Senior/Staff DevOps engineer building automated testing and delivery infrastructure for TurbineOne’s military edge systems. Improving platform reliability, developer productivity, and software quality across remote hardware.
🇺🇸 United States – Remote
💵 $240k - $280k / year
💰 $3M Seed Round on 2022-01
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 2 hours ago
Lead DevOps Engineer driving AWS migration and modernization for EverOps, an embedded infrastructure partner. Building networking, deployment, observability, and disaster-recovery platforms for payments environments.
🇺🇸 United States – Remote
💰 Pre Seed Round on 2022-02
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 3 hours ago
Quality and reliability engineer qualifying AI server power modules at Renesas, a global semiconductor solutions company. Developing reliability tests, component qualification processes, validation plans, and manufacturing prototypes for high-performance computing products.
🇺🇸 United States – Remote
💰 Venture Round on 2022-06
⏰ Full Time
🟠 Senior
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🔥 4 hours ago
Manager de banco de dados liderando disponibilidade, performance e segurança para a Stone, empresa brasileira de tecnologia e serviços financeiros. Condução de equipes, cloud migration e confiabilidade de plataformas críticas.
🇺🇸 United States – Remote
⏰ Full Time
🟢 Junior
🟡 Mid-level
⛑ DevOps & Site Reliability Engineer (SRE)
🚫👨🎓 No degree required
🦅 H1B Visa Sponsor
🗣️🇧🇷🇵🇹 Portuguese Required
🔥 6 hours ago
Senior SRE operating Datavant’s healthcare data and ML platform remotely in the United States. Building reliable cloud infrastructure, observability, CI/CD, and cross-platform data systems.
🇺🇸 United States – Remote
💵 $168k - $200k / year
💰 $40M Series B on 2020-10
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor