K8 Site Reliability SME

🕒 Agosto 13

🏄 California, Texas – Remoto

infoinfo

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 18%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 funcionários

💼 Consultoria

📦 Logística

🏗️ Construção

💰 Post-IPO Equity em 2023-05

Consulting • Logistics • Construction

A Bitdeer Technologies Group (Nasdaq: BTDR) é uma líder na indústria de blockchain e computação de alto desempenho. É uma das maiores detentoras mundiais de hash rate proprietário e fornecedoras de hash rate. A Bitdeer está comprometida em fornecer soluções de computação abrangentes para seus clientes.

Descrição

• Design, deploy, and operate production Kubernetes control planes optimized for GPU workloads at scale (100–10,000 GPUs) • Configure Nvidia GPU Operator, device plugin, MIG, and GPU time-slicing policies • Implement topology-aware scheduling using GPU locality, NVLink domain awareness, and network rail affinity • Build Custom Resource Definitions (CRDs) for GPU workload lifecycle management • Integrate Slurm on K8S, Ray on K8S, and Kubeflow • Enforce multi-tenant isolation through namespaces, network policies, resource quotas, RBAC, and pod security standards • Automate Bare-Metal-as-a-Service provisioning, tenant onboarding, lifecycle management, and reclamation • Develop Terraform providers and modules for infrastructure-as-code across GPU clusters • Define and publish SLIs/SLOs for cluster availability, job completion rates, and provisioning latency • Automate incident-management runbooks, escalation, and post-incident reviews • Operate the Prometheus, Grafana, Alertmanager, and PagerDuty monitoring stack • Automate GPU node failure detection, drain/cordon/taint, and workload rescheduling • Make the Kubernetes control plane safe for automated AIOps remediation • Provide CRD schemas for platform predictors and remediators • Turn human interventions into autonomous workflows • Deliver automated drain/reschedule around predicted GPU faults without customer impact • Launch BMaaS for external tenants with self-service onboarding • Meet cluster-availability and job-completion SLOs

🎯 Requisitos

• 5+ years in Kubernetes operations, including at least 2 years managing GPU workloads on K8S • Deep understanding of Nvidia GPU Operator, device plugin, and GPU scheduling in K8S • Experience with topology-aware scheduling and GPU-specific resource management • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees • Experience with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom solutions • Proficiency in Terraform, Helm, and GitOps workflows such as ArgoCD or Flux • Strong SRE background, including SLI/SLO frameworks, incident management, and capacity planning • Experience with Prometheus, Grafana, and alerting at scale • Strong programming skills in Go or Python for operator/CRD development • AIOps aptitude: ability to treat the K8S control plane as an execution surface for automated remediation • Experience wiring an autoscaler/remediator loop into K8S, or ability to design one • Runbook-as-code mindset • Equal employment opportunity compliance with applicable country, state, and local laws

Candidatar-se

Vagas Similares

🕒 Agosto 13

TurbineOne

11 - 50

🎖️ Defesa

🏭 Manufatura

🚀 Aeroespacial

Senior/Staff DevOps engineer building automated testing and delivery infrastructure for TurbineOne’s military edge systems. Improving platform reliability, developer productivity, and software quality across remote hardware.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $240.000 - $280.000 / ano

💰 $3.000.000 Seed Round em 2022-01

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 13

Datavant

201 - 500

🏥 Saúde

💼 Consultoria

⚕️ Seguro de Saúde

Senior SRE operating Datavant’s healthcare data and ML platform remotely in the United States. Building reliable cloud infrastructure, observability, CI/CD, and cross-platform data systems.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $168.000 - $200.000 / ano

💰 $40.000.000 Series B em 2020-10

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 13

Sumsub

501 - 1000

💼 Consultoria

⚖️ Jurídico

📦 Logística

DevSecOps Engineer building automated vulnerability and fleet security capabilities for Sumsub’s AI-powered trust infrastructure. Creating self-service APIs, dashboards, and compliance controls for global engineering teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $30.000.000 Series B - Sumsub em 2022-12

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 13

Union Home Mortgage Corp.

1001 - 5000

🏦 Bancário

🏠 Imobiliário

Infrastructure DevOps Engineer modernizing Union Home Mortgage’s AWS and Azure cloud platforms. Building automation, CI/CD pipelines, infrastructure as code, and operational processes for enterprise systems.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 12

Mirantis

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

Senior DevOps Engineer operating high-performance Kubernetes storage for Mirantis, a Kubernetes-native AI infrastructure company. Automating NFS, Linux, and air-gapped storage platforms for GPU workloads at scale.

🗣️🇺🇸🇬🇧 Inglês obrigatório