Senior GPU Cloud, K8S Expert

Vaga não está no LinkedIn

🕒 Setembro 11

🏄 California, Texas – Remoto

infoinfo

💵 $180.000 - $260.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 8%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 funcionários

💼 Consultoria

📦 Logística

🏗️ Construção

💰 Post-IPO Equity em 2023-05

Consulting • Logistics • Construction

A Bitdeer Technologies Group (Nasdaq: BTDR) é uma líder na indústria de blockchain e computação de alto desempenho. É uma das maiores detentoras mundiais de hash rate proprietário e fornecedoras de hash rate. A Bitdeer está comprometida em fornecer soluções de computação abrangentes para seus clientes.

Descrição

• Design, deploy, and operate the control plane for an AI-operated GPU cloud • Run production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs) • Configure Nvidia GPU operator, device plugin, MIG, and GPU time-slicing policies • Implement topology-aware scheduling using GPU locality, NVLink domain awareness, and network rail affinity • Develop Custom Resource Definitions (CRDs) for GPU workload lifecycle management • Integrate Slurm, Ray, and Kubeflow with Kubernetes • Enforce multi-tenant isolation through namespaces, network policies, resource quotas, RBAC, and pod security standards • Automate BMaaS provisioning, tenant onboarding, lifecycle, and reclamation • Build Terraform providers and modules for infrastructure-as-code across GPU clusters • Define and meet SLIs/SLOs for cluster availability, job completion rates, and provisioning latency • Automate incident management, escalation, and post-incident reviews • Operate monitoring with Prometheus, Grafana, Alertmanager, and PagerDuty • Automate GPU node failure detection, drain/cordon/taint, and workload rescheduling • Make the control plane safe for automated AIOps remediation • Enable automated workflows for predicted GPU faults without customer impact • Launch BMaaS for external tenants with self-service onboarding • Publish and meet cluster availability and job-completion SLOs

🎯 Requisitos

• 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S • Experience with topology-aware scheduling and GPU-specific resource management • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom) • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux) • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning • Experience with Prometheus, Grafana, and alerting at scale • Strong programming skills in Go or Python for operator/CRD development • AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler • You've either wired an autoscaler/remediator loop into K8S or you can design one • Runbook-as-code mindset — every SRE playbook you write should be executable by the platform

Candidatar-se

Vagas Similares

🕒 Setembro 11

Paylocity

5001 - 10000

👥 RH Tech

☁️ SaaS

🤝 B2B

DevSecOps Engineer securing Paylocity’s cloud-based HR and payroll software platform. Developing security tooling, integrating build protections, and guiding vulnerability remediation across web and mobile applications.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $96.000 - $130.000 / ano

💰 $10.000.000 Venture Round - Paylocity em 2008-05

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 11

Ookla

201 - 500

📡 Telecomunicações

🏢 Corporativo

Site Reliability Engineer maintaining Ookla’s global cloud, database, and observability infrastructure. Supporting connectivity intelligence services used by hundreds of millions worldwide.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $90.000 - $100.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 10

Zafran Security

51 - 200

🔐 Segurança

Senior DevOps Engineer owning Zafran’s US cybersecurity production environment. Building AWS infrastructure, CI/CD, Kubernetes operations, and production reliability.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 10

Delinea

1001 - 5000

🔒 Cibersegurança

☁️ SaaS

🏢 Corporativo

Senior Site Reliability Engineer owning reliability, monitoring, and incident response for Delinea’s FedRAMP SaaS identity security platform. Automating Azure and AWS operations.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $130.000 - $160.000 / ano

💰 Private Equity Round em 2021-03

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 10

Avanade

10.000+ funcionários

💼 Consultoria

📦 Logística

📣 Marketing

Avanade manager architecting Azure DevOps, GitHub, and AI-enabled software delivery solutions for enterprise clients. Leading DevOps transformation, Copilot adoption, governance, and cloud engineering modernization.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $139.200 - $195.700 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório