Principal Operations Engineer – Reliability, Data Center Operations

🕒 Junho 29

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $250.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of FluidStack

FluidStack

11 - 50 funcionários

🤖 Inteligência Artificial

Artificial Intelligence • Cloud Computing

A FluidStack é uma empresa que fornece infraestrutura de supercomputação com GPUs para laboratórios de IA. Ela oferece acesso sob demanda a milhares de GPUs Nvidia, permitindo treinamento e inferência de IA em larga escala. A empresa se especializa em implantar e gerenciar grandes clusters de GPUs, com suporte para tecnologias como Kubernetes e Slurm, garantindo alta disponibilidade e excelente suporte. A FluidStack fornece uma infraestrutura de nuvem totalmente gerenciada, ajudando empresas de IA a focarem no desenvolvimento de modelos sem se preocuparem com o hardware subjacente. Eles enfatizam desempenho e eficiência de custos, oferecendo serviços que escalam para milhares de GPUs com alta disponibilidade e tempos de resposta rápidos.

Descrição

• Take the on-call escalation when a site hits trouble and triage it virtually, using real knowledge of the team and the systems to decide what to escalate, when, and how to keep the field crew focused without burying them. • Get on a plane when it matters: travel site to site (50%+) to work live incidents and post-incident reviews on the floor, and bring the practices that worked elsewhere with you. • Own root cause analysis on significant events through to closure and track corrective actions to done, killing the underlying class of failure rather than the one instance in front of you. • Read the patterns across the fleet’s incidents and RCAs, push the few highest-value learnings through to closure, and stay honest about what’s achievable and what to drop instead of boiling the ocean. • Carry learnings and practices from one campus to the next so a fix at one site becomes the standard everywhere before the failure repeats. • Write the operational Assessment standard and audit each campus against it, feeding what you find straight back into the corrective-action loop.

🎯 Requisitos

• You’ve run a live critical operation and led a team of operators, and you carry the deep, earned judgment that comes from owning the floor when it counts. • You’ve been the person a site calls when something breaks, triaged the problem over the phone, and known exactly when to escalate and when to let the field team work it. • You’ve authored root cause analyses on significant events and tracked corrective actions to closure, and you can show the difference between an RCA that closed a ticket and one that killed a class of failure. • You’ve sat with a pile of RCA actions and cut it to the few that matter, because you know an operation that commits to everything finishes nothing. • You’ve traveled site to site, walked the floor, and left each operation better than you found it, carrying the practices that worked from one into the next. • You’ve written the standard, not just followed it, audited real sites against it without flinching from what you found, and can hold one bar across domains you don’t all live in. • Bonus: Hyperscale or large colocation at hundreds of MW+. Direct exposure to Hardware or Network operations, not only Facilities, incident.io or equivalent incident tooling, plus DCIM. Building an assessment, audit, qualification, or training program from scratch.

🏖️ Benefícios

• Competitive total compensation package (salary + equity). • Retirement or pension plan, in line with local norms. • Health, dental, and vision insurance. • Generous PTO policy, in line with local norms.

Candidatar-se

Vagas Similares

🕒 Junho 24

Redox

201 - 500

🏥 Saúde

⚕️ Seguro de Saúde

☁️ SaaS

DevSecOps Engineer ensuring secure software development at Redox, enhancing healthcare data exchange. Collaborating with platform engineers to implement security best practices across the AWS/EKS infrastructure.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 23

Lyric - Clarity in motion.

201 - 500

🏥 Saúde

💼 Consultoria

📦 Logística

Azure DevOps Engineer at Lyric managing Azure infrastructure for healthcare technology solutions. Focus on security, reliability, and operational efficiency in a remote role.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.289 - $225.434 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 23

SAIC

10.000+ funcionários

☁️ SaaS

📣 Marketing

🏢 Corporativo

DevSecOps Engineer providing exceptional DevOps engineering for advancing CI/CD and automating pipelines. Must have deep proficiency in AWS, Azure, and DevSecOps tools.

🇺🇸 Estados Unidos – Remoto (EUA)

🔥 Investimento no último ano

💰 $500.000.000 Post-IPO Debt - SAIC em 2025-09

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 22

Kong Inc.

201 - 500

💼 Consultoria

📦 Logística

🔌 API

Staff Site Reliability Engineer for Kong's Volcano platform overseeing reliability and infrastructure scaling. Collaborating on SRE practices and emerging technology evaluations.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $210.000 / ano

💰 $100.000.000 Series D em 2021-02

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 20

Gorilla Logic

501 - 1000

💼 Consultoria

📣 Marketing

📦 Logística

Technical Engineering Manager leading high-performing cloud and DevOps teams. Guiding architecture and delivery of scalable, reliable, and secure cloud solutions for clients.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório