System Reliability Engineering Lead

🕒 5 dias atrás

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $151.800 - $227.700 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 0%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of GE Vernova

GE Vernova

10.000+ funcionários

💼 Consultoria

📦 Logística

🏭 Manufatura

Consulting • Logistics • Manufacturing

GE Vernova é líder no setor de energia, com mais de 130 anos de experiência, dedicada a eletrificar o mundo enquanto o descarboniza. A empresa oferece um amplo portfólio de soluções em energia, incluindo tecnologias de geração a gás, hidrelétrica, nuclear e eólica, com o objetivo de fornecer energia confiável, acessível e sustentável. Com forte foco em inovação, a GE Vernova desempenha um papel significativo na redução da pegada de carbono dos sistemas de energia globais e apoia a transição para emissões líquidas zero (net zero) até 2030.

Descrição

• Serve as the hands-on technical authority for production stability across the GridOS SaaS portfolio • Own Change Management and approve or halt production deployments based on system health • Drive high reliability and engineering excellence across a distributed team • Architect and implement standardized, secure cloud infrastructure provisioning • Automate account provisioning to accelerate customer onboarding • Define and build the standardized Middle-Mile software delivery platform using Backstage, ArgoCD, and GitHub Actions • Establish global handover protocols and 24/7 operational coverage across US, India, and Mexico time zones • Establish and own enterprise-wide SLOs, SLIs, and error budgets • Serve as final technical authority for production releases and enforce security and performance quality gates • Implement Canary and Blue/Green deployments with automated rollback capabilities • Build and mature the SRE Center for Enablement with coaching, templates, and reliability patterns • Lead incident response for Sev1/Sev2 events and P1 escalations • Facilitate blameless Root Cause Analysis and own the post-incident lifecycle • Architect and validate backup and disaster recovery strategies, including cross-region failover and automated recovery testing • Own FinOps, cloud cost optimization, and long-term capacity planning • Serve as primary SRE point of contact for North American utility customers • Participate in customer reviews, incident communications, and service health reporting • Lead a distributed team of 8 SRE engineers across Hyderabad and Querétaro • Set technical direction, assign tasks, own deliverables, mentor engineers, and provide performance feedback to the people leader of record • Travel up to 10% to customer sites and team locations as needed

🎯 Requisitos

• Deep expertise in AWS core services: EC2, EKS, RDS, S3, and IAM • Experience with AWS management tools including CloudTrail and CloudWatch • Advanced mastery of Kubernetes internals and EKS cluster operations across multi-region architectures • Expert knowledge of ArgoCD, GitHub Actions, and GitOps-first workflows • Proficiency in Infrastructure as Code using Terraform • Proficiency in configuration management via Ansible • Hands-on experience with Prometheus, Grafana, Splunk or Datadog, and OpenTelemetry • Experience with cloud cost optimization, reserved instance management, right-sizing, and long-term capacity planning for multi-tenant SaaS platforms • 12+ years in software engineering, cloud operations, or infrastructure roles • 8–10 years of hands-on experience in SRE, Platform Engineering, Cloud Operations, or Production Support for large-scale, distributed SaaS applications • Proven track record leading distributed engineering teams as a player-coach while remaining hands-on with architecture, automation, and incident response • Exceptional troubleshooting skills under pressure and a “Fire Marshal” mindset toward investigation and proactive inspection • Experience working directly with enterprise customers on production reliability, incident communication, and service-level reporting • Must pass customer-mandated background screening for access to critical infrastructure environments • Must be legally authorized to work in the United States • Must complete a drug screen, as applicable • General shift during US business hours and on-call availability for P1/Sev1 incidents • Up to 10% travel to customer sites and team locations • Desired: NERC CIP, SOC2, ISO 27001, or IEC 62443 knowledge/experience • Desired: experience in highly regulated industries such as utilities, financial services, or critical national infrastructure • Desired certifications: AWS DevOps Engineer—Professional or Solutions Architect—Associate/Professional, CKA, SRE Practitioner, and AWS FinOps Practitioner or equivalent

🏖️ Benefícios

• Discretionary annual bonus • Medical, dental, vision, and prescription drug coverage • Health Coach from GE Vernova, a 24/7 nurse-based resource • Employee Assistance Program with 24/7 confidential assessment, counseling, and referral services • GE Vernova Retirement Savings Plan • Tax-advantaged 401(k) savings opportunity with company matching contributions and company retirement contributions • Fidelity resources and financial planning consultants • Tuition assistance • Adoption assistance • Paid parental leave • Disability benefits • Life insurance • 12 paid holidays • Permissive time off • Professional development opportunities • Relocation assistance not provided

Candidatar-se

Vagas Similares

🕒 5 dias atrás

Millennium

201 - 500

💼 Consultoria

🎖️ Defesa

🔒 Cibersegurança

Remote DevSecOps cybersecurity engineer securing DoD software, GitLab CI/CD pipelines, containers, and vulnerability management. Supporting Millennium’s national-security cybersecurity missions through RMF compliance and software assurance.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $125.000 - $135.000 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 5 dias atrás

GitLab

1001 - 5000

💼 Consultoria

📣 Marketing

🤖 Inteligência Artificial

Distinguished Engineer directing GitLab’s CI, CD, Plan, and source-control architecture. Driving scalable AI-native DevOps systems, technical strategy, and engineering quality.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $250.000 - $349.000 / ano

💰 Secondary Market em 2020-11

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 5 dias atrás

Vultr

201 - 500

🤖 Inteligência Artificial

🤝 B2B

🔧 Hardware

Senior SRE maintaining MySQL and PostgreSQL reliability for Vultr’s global cloud infrastructure. Owning monitoring, disaster recovery, incident response, security compliance, and automation.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $125.000 - $135.000 / ano

💰 $329.000.000 Debt Financing - Vultr em 2025-06

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 5 dias atrás

CACI International Inc

10.000+ funcionários

🎖️ Defesa

🏛️ Governo

🔒 Cibersegurança

Operating System Deployment Engineer managing secure Windows and Windows Server images for CACI’s DoD enterprise IT services. Automating deployments and supporting physical and virtual infrastructure across 187 bases.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $75.200 - $158.100 / ano

🔥 Investimento no último ano

💰 $500.000.000 Post-IPO Debt em 2026-02

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 5 dias atrás

PingWind Inc. (SDVOSB)

51 - 200

💼 Consultoria

📦 Logística

🏥 Saúde

DevSecOps Engineer securing cloud-native software delivery for PingWind, a federal government services provider. Building CI/CD security, automating controls, and supporting vulnerability remediation.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório