
201 - 500 employees
💼 Consulting
📦 Logistics
🏗️ Construction
💰 Post-IPO Equity on 2023-05
Consulting • Logistics • Construction
Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.
🕒 August 13
🏄 California, Texas – Remote
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
👻 Ghost score 14%
Improve your chances of getting an interview by checking your resume score before you apply.

201 - 500 employees
💼 Consulting
📦 Logistics
🏗️ Construction
💰 Post-IPO Equity on 2023-05
Consulting • Logistics • Construction
Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.
• Design, deploy, and operate production Kubernetes control planes optimized for GPU workloads at scale (100–10,000 GPUs) • Configure Nvidia GPU Operator, device plugin, MIG, and GPU time-slicing policies • Implement topology-aware scheduling using GPU locality, NVLink domain awareness, and network rail affinity • Build Custom Resource Definitions (CRDs) for GPU workload lifecycle management • Integrate Slurm on K8S, Ray on K8S, and Kubeflow • Enforce multi-tenant isolation through namespaces, network policies, resource quotas, RBAC, and pod security standards • Automate Bare-Metal-as-a-Service provisioning, tenant onboarding, lifecycle management, and reclamation • Develop Terraform providers and modules for infrastructure-as-code across GPU clusters • Define and publish SLIs/SLOs for cluster availability, job completion rates, and provisioning latency • Automate incident-management runbooks, escalation, and post-incident reviews • Operate the Prometheus, Grafana, Alertmanager, and PagerDuty monitoring stack • Automate GPU node failure detection, drain/cordon/taint, and workload rescheduling • Make the Kubernetes control plane safe for automated AIOps remediation • Provide CRD schemas for platform predictors and remediators • Turn human interventions into autonomous workflows • Deliver automated drain/reschedule around predicted GPU faults without customer impact • Launch BMaaS for external tenants with self-service onboarding • Meet cluster-availability and job-completion SLOs
• 5+ years in Kubernetes operations, including at least 2 years managing GPU workloads on K8S • Deep understanding of Nvidia GPU Operator, device plugin, and GPU scheduling in K8S • Experience with topology-aware scheduling and GPU-specific resource management • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees • Experience with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom solutions • Proficiency in Terraform, Helm, and GitOps workflows such as ArgoCD or Flux • Strong SRE background, including SLI/SLO frameworks, incident management, and capacity planning • Experience with Prometheus, Grafana, and alerting at scale • Strong programming skills in Go or Python for operator/CRD development • AIOps aptitude: ability to treat the K8S control plane as an execution surface for automated remediation • Experience wiring an autoscaler/remediator loop into K8S, or ability to design one • Runbook-as-code mindset • Equal employment opportunity compliance with applicable country, state, and local laws
Apply Now🕒 August 13
Senior/Staff DevOps engineer building automated testing and delivery infrastructure for TurbineOne’s military edge systems. Improving platform reliability, developer productivity, and software quality across remote hardware.
🇺🇸 United States – Remote
💵 $240k - $280k / year
💰 $3M Seed Round on 2022-01
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 August 13
Manager de banco de dados liderando disponibilidade, performance e segurança para a Stone, empresa brasileira de tecnologia e serviços financeiros. Condução de equipes, cloud migration e confiabilidade de plataformas críticas.
🇺🇸 United States – Remote
⏰ Full Time
🟢 Junior
🟡 Mid-level
⛑ DevOps & Site Reliability Engineer (SRE)
🚫👨🎓 No degree required
🦅 H1B Visa Sponsor
🗣️🇧🇷🇵🇹 Portuguese Required
🕒 August 13
Senior SRE operating Datavant’s healthcare data and ML platform remotely in the United States. Building reliable cloud infrastructure, observability, CI/CD, and cross-platform data systems.
🇺🇸 United States – Remote
💵 $168k - $200k / year
💰 $40M Series B on 2020-10
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🕒 August 13
DevSecOps Engineer building automated vulnerability and fleet security capabilities for Sumsub’s AI-powered trust infrastructure. Creating self-service APIs, dashboards, and compliance controls for global engineering teams.
🇺🇸 United States – Remote
💰 $30M Series B - Sumsub on 2022-12
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 August 13
Infrastructure DevOps Engineer modernizing Union Home Mortgage’s AWS and Azure cloud platforms. Building automation, CI/CD pipelines, infrastructure as code, and operational processes for enterprise systems.