
201 - 500 employees
đź Consulting
đŚ Logistics
đď¸ Construction
đ° Post-IPO Equity on 2023-05
Consulting ⢠Logistics ⢠Construction
Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the worldâs largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.
đĽ 12 hours ago
đ California, Texas â Remote
đľ $180k - $260k / year
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đť Ghost score 5%
Ansible
AWS
Azure
Cloud
Distributed Systems
DNS
Docker
Google Cloud Platform
Grafana
Kubernetes
Linux
Prometheus
Python
TCP/IP
Terraform
Go
Improve your chances of getting an interview by checking your resume score before you apply.

201 - 500 employees
đź Consulting
đŚ Logistics
đď¸ Construction
đ° Post-IPO Equity on 2023-05
Consulting ⢠Logistics ⢠Construction
Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the worldâs largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.
⢠Design, implement, and maintain end-to-end CI/CD pipelines for software applications and machine learning models ⢠Automate build, test, deployment, and rollback processes ⢠Build, optimize, and scale cloud-native infrastructure using Kubernetes and Docker ⢠Manage and provision specialized computing resources, including GPU clusters, for high-performance AI workloads and model inferencing ⢠Own high-availability design in production environments ⢠Implement disaster recovery strategies, self-healing mechanisms, capacity planning, and performance tuning ⢠Champion Infrastructure as Code practices using Terraform, Ansible, and Helm ⢠Architect and refine monitoring, logging, and alerting systems using Prometheus, Grafana, and ELK/EFK stack ⢠Build the Internal Developer Platform and golden paths enabling product, model, and data-science teams to deploy without opening a ticket ⢠Collaborate with R&D, Data Science, Security, and Business teams to streamline workflows and eliminate bottlenecks ⢠Establish and enforce system stability and security standards ⢠Manage release workflows, implement Zero Trust access controls, oversee secrets management, and ensure SOC2/ISO27001 compliance ⢠Lead troubleshooting, root-cause analysis, and preventative remediation during complex anomalies and major incidents ⢠Convert incident learnings into automation to prevent recurrence
⢠Bachelor's degree or above in Computer Science, Engineering, or a related technical field ⢠5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles ⢠Expert-level knowledge of Linux operating systems and core networking principles, including TCP/IP, DNS, HTTP, Load Balancing, and VPCs ⢠Deep mastery of Docker and Kubernetes orchestration, including cluster management and production-level best practices ⢠Proficiency designing and managing infrastructure on major public or hybrid cloud platforms, including AWS, GCP, Azure, or Alibaba Cloud ⢠Experience with multi-cloud and hybrid-cloud strategies ⢠Strong coding and scripting capabilities in at least one major language such as Go, Python, or Shell ⢠Systematic and practical understanding of CI/CD methodologies, Infrastructure as Code (IaC), observability paradigms, and SRE principles ⢠Exceptional problem-solving abilities and sharp technical judgment ⢠Excellent cross-team communication skills ⢠Preferred: Familiarity with MLOps practices, model serving/inferencing frameworks such as vLLM, TGI, or Triton Inference Server ⢠Preferred: Experience managing GPU clusters for AI/ML workloads ⢠Preferred: Experience with large-scale distributed systems or high-concurrency environments ⢠Preferred: Hands-on experience designing and building Internal Developer Platforms (IDP) ⢠Preferred: Familiarity with Zero Trust architecture, automated security testing (DevSecOps), SOC2, or ISO27001 ⢠Preferred: Prior experience as a Technical Lead, mentoring junior engineers, or managing DevOps teams ⢠Preferred: Experience wiring an LLM-driven code/config helper into a pipeline or strong opinions on how to ⢠Must comply with applicable work authorization and equal employment requirements in the relevant country, state, and local jurisdictions
Apply NowđĽ 12 hours ago
Senior SRE operating Bitwarden Govâs FedRAMP-compliant cloud infrastructure. Managing reliability, monitoring, incident response, Kubernetes, and security across multi-cloud environments.
đşđ¸ United States â Remote
đľ $140k - $185k / year
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đĽ 12 hours ago
DevOps Team Lead operating VASâs AWS platform for farm management software. Leading IaC, reliability, security, cost optimization, and globally distributed workloads.
đĽ 14 hours ago
Release Train Engineer leading SAFe delivery and DevOps systems for VetsEZâs federal healthcare IT project. Overseeing Agile Release Train execution, CI/CD, platform operations, and cross-team delivery.
đĽ 19 hours ago
Site Reliability Engineer operating Azure infrastructure for System Automationâs regulatory-agency SaaS platform. Automating reliability, observability, CI/CD, security, and incident response.
đşđ¸ United States â Remote
đľ $120k - $140k / year
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đ Yesterday
Senior DevOps Engineer strengthening Worth AIâs cloud infrastructure, Kubernetes, and reliability systems. Automating AWS infrastructure, CI/CD, observability, disaster recovery, and platform resilience.