Senior GPU Cloud, K8S Expert

Job not on LinkedIn

🔥 12 minutes ago

🏄 California, Texas – Remote

infoinfo

💵 $180k - $260k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 4%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 employees

💼 Consulting

📦 Logistics

🏗️ Construction

💰 Post-IPO Equity on 2023-05

Consulting • Logistics • Construction

Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.

📋 Description

• Design, deploy, and operate the control plane for an AI-operated GPU cloud • Run production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs) • Configure Nvidia GPU operator, device plugin, MIG, and GPU time-slicing policies • Implement topology-aware scheduling using GPU locality, NVLink domain awareness, and network rail affinity • Develop Custom Resource Definitions (CRDs) for GPU workload lifecycle management • Integrate Slurm, Ray, and Kubeflow with Kubernetes • Enforce multi-tenant isolation through namespaces, network policies, resource quotas, RBAC, and pod security standards • Automate BMaaS provisioning, tenant onboarding, lifecycle, and reclamation • Build Terraform providers and modules for infrastructure-as-code across GPU clusters • Define and meet SLIs/SLOs for cluster availability, job completion rates, and provisioning latency • Automate incident management, escalation, and post-incident reviews • Operate monitoring with Prometheus, Grafana, Alertmanager, and PagerDuty • Automate GPU node failure detection, drain/cordon/taint, and workload rescheduling • Make the control plane safe for automated AIOps remediation • Enable automated workflows for predicted GPU faults without customer impact • Launch BMaaS for external tenants with self-service onboarding • Publish and meet cluster availability and job-completion SLOs

🎯 Requirements

• 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S • Experience with topology-aware scheduling and GPU-specific resource management • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom) • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux) • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning • Experience with Prometheus, Grafana, and alerting at scale • Strong programming skills in Go or Python for operator/CRD development • AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler • You've either wired an autoscaler/remediator loop into K8S or you can design one • Runbook-as-code mindset — every SRE playbook you write should be executable by the platform

Apply Now

Similar Jobs

🔥 2 hours ago

Centex Technologies

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

DevOps Engineer 4 building secure CI/CD and cloud platforms for Centex Technologies’ programs and customers. Automating infrastructure, observability, security, and reliable software delivery.

🔥 3 hours ago

Paylocity

5001 - 10000

👥 HR Tech

☁️ SaaS

🤝 B2B

DevSecOps Engineer securing Paylocity’s cloud-based HR and payroll software platform. Developing security tooling, integrating build protections, and guiding vulnerability remediation across web and mobile applications.

🇺🇸 United States – Remote

💵 $96k - $130k / year

💰 $10M Venture Round - Paylocity on 2008-05

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 6 hours ago

EIS Ltd

1001 - 5000

💳 Fintech

🤝 B2B

🛡️ Insurance

Senior DevOps Engineer owning AWS infrastructure, Kubernetes, CI/CD, and observability for EIS SaaS applications. Building secure, reliable end-to-end delivery platforms.

🔥 14 hours ago

Ookla

201 - 500

📡 Telecommunications

🏢 Enterprise

Site Reliability Engineer maintaining Ookla’s global cloud, database, and observability infrastructure. Supporting connectivity intelligence services used by hundreds of millions worldwide.

🔥 17 hours ago

Zafran Security

51 - 200

🔐 Security

Senior DevOps Engineer owning Zafran’s US cybersecurity production environment. Building AWS infrastructure, CI/CD, Kubernetes operations, and production reliability.