K8 Site Reliability SME

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 employees

💼 Consulting

📦 Logistics

🏗️ Construction

💰 Post-IPO Equity on 2023-05

Consulting • Logistics • Construction

Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.

📋 Description

• Design, deploy, and operate production Kubernetes control planes optimized for GPU workloads at scale (100–10,000 GPUs) • Configure Nvidia GPU Operator, device plugin, MIG, and GPU time-slicing policies • Implement topology-aware scheduling using GPU locality, NVLink domain awareness, and network rail affinity • Build Custom Resource Definitions (CRDs) for GPU workload lifecycle management • Integrate Slurm on K8S, Ray on K8S, and Kubeflow • Enforce multi-tenant isolation through namespaces, network policies, resource quotas, RBAC, and pod security standards • Automate Bare-Metal-as-a-Service provisioning, tenant onboarding, lifecycle management, and reclamation • Develop Terraform providers and modules for infrastructure-as-code across GPU clusters • Define and publish SLIs/SLOs for cluster availability, job completion rates, and provisioning latency • Automate incident-management runbooks, escalation, and post-incident reviews • Operate the Prometheus, Grafana, Alertmanager, and PagerDuty monitoring stack • Automate GPU node failure detection, drain/cordon/taint, and workload rescheduling • Make the Kubernetes control plane safe for automated AIOps remediation • Provide CRD schemas for platform predictors and remediators • Turn human interventions into autonomous workflows • Deliver automated drain/reschedule around predicted GPU faults without customer impact • Launch BMaaS for external tenants with self-service onboarding • Meet cluster-availability and job-completion SLOs

🎯 Requirements

• 5+ years in Kubernetes operations, including at least 2 years managing GPU workloads on K8S • Deep understanding of Nvidia GPU Operator, device plugin, and GPU scheduling in K8S • Experience with topology-aware scheduling and GPU-specific resource management • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees • Experience with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom solutions • Proficiency in Terraform, Helm, and GitOps workflows such as ArgoCD or Flux • Strong SRE background, including SLI/SLO frameworks, incident management, and capacity planning • Experience with Prometheus, Grafana, and alerting at scale • Strong programming skills in Go or Python for operator/CRD development • AIOps aptitude: ability to treat the K8S control plane as an execution surface for automated remediation • Experience wiring an autoscaler/remediator loop into K8S, or ability to design one • Runbook-as-code mindset • Equal employment opportunity compliance with applicable country, state, and local laws

Apply Now

Similar Jobs

🔥 34 minutes ago

TurbineOne

11 - 50

🎖️ Defense

🏭 Manufacturing

🚀 Aerospace

Senior/Staff DevOps engineer building automated testing and delivery infrastructure for TurbineOne’s military edge systems. Improving platform reliability, developer productivity, and software quality across remote hardware.

🔥 2 hours ago

EverOps

51 - 200

🤝 B2B

🏢 Enterprise

☁️ SaaS

Lead DevOps Engineer driving AWS migration and modernization for EverOps, an embedded infrastructure partner. Building networking, deployment, observability, and disaster-recovery platforms for payments environments.

🔥 3 hours ago

Renesas Electronics

10,000+ employees

🏭 Manufacturing

🏥 Healthcare

📦 Logistics

Quality and reliability engineer qualifying AI server power modules at Renesas, a global semiconductor solutions company. Developing reliability tests, component qualification processes, validation plans, and manufacturing prototypes for high-performance computing products.

🔥 4 hours ago

Stone & Company

1 - 10

💼 Consulting

🤝 B2B

Manager de banco de dados liderando disponibilidade, performance e segurança para a Stone, empresa brasileira de tecnologia e serviços financeiros. Condução de equipes, cloud migration e confiabilidade de plataformas críticas.

🗣️🇧🇷🇵🇹 Portuguese Required

🔥 6 hours ago

Datavant

201 - 500

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Senior SRE operating Datavant’s healthcare data and ML platform remotely in the United States. Building reliable cloud infrastructure, observability, CI/CD, and cross-platform data systems.