System Reliability Engineering Lead

🔥 21 hours ago

🇺🇸 United States – Remote

💵 $151.8k - $227.7k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of GE Vernova

GE Vernova

10,000+ employees

💼 Consulting

📦 Logistics

🏭 Manufacturing

Consulting • Logistics • Manufacturing

GE Vernova is a leader in the energy sector with over 130 years of experience, dedicated to electrifying the world while decarbonizing it. The company offers a broad portfolio of energy solutions including gas, hydro, nuclear, and wind power technologies, aimed at providing reliable, affordable, and sustainable energy. With a strong focus on innovation, GE Vernova plays a significant role in reducing the carbon footprint of global power systems and supports the transition to net-zero emissions by 2030.

📋 Description

• Serve as the hands-on technical authority for production stability across the GridOS SaaS portfolio • Own Change Management and approve or halt production deployments based on system health • Drive high reliability and engineering excellence across a distributed team • Architect and implement standardized, secure cloud infrastructure provisioning • Automate account provisioning to accelerate customer onboarding • Define and build the standardized Middle-Mile software delivery platform using Backstage, ArgoCD, and GitHub Actions • Establish global handover protocols and 24/7 operational coverage across US, India, and Mexico time zones • Establish and own enterprise-wide SLOs, SLIs, and error budgets • Serve as final technical authority for production releases and enforce security and performance quality gates • Implement Canary and Blue/Green deployments with automated rollback capabilities • Build and mature the SRE Center for Enablement with coaching, templates, and reliability patterns • Lead incident response for Sev1/Sev2 events and P1 escalations • Facilitate blameless Root Cause Analysis and own the post-incident lifecycle • Architect and validate backup and disaster recovery strategies, including cross-region failover and automated recovery testing • Own FinOps, cloud cost optimization, and long-term capacity planning • Serve as primary SRE point of contact for North American utility customers • Participate in customer reviews, incident communications, and service health reporting • Lead a distributed team of 8 SRE engineers across Hyderabad and Querétaro • Set technical direction, assign tasks, own deliverables, mentor engineers, and provide performance feedback to the people leader of record • Travel up to 10% to customer sites and team locations as needed

🎯 Requirements

• Deep expertise in AWS core services: EC2, EKS, RDS, S3, and IAM • Experience with AWS management tools including CloudTrail and CloudWatch • Advanced mastery of Kubernetes internals and EKS cluster operations across multi-region architectures • Expert knowledge of ArgoCD, GitHub Actions, and GitOps-first workflows • Proficiency in Infrastructure as Code using Terraform • Proficiency in configuration management via Ansible • Hands-on experience with Prometheus, Grafana, Splunk or Datadog, and OpenTelemetry • Experience with cloud cost optimization, reserved instance management, right-sizing, and long-term capacity planning for multi-tenant SaaS platforms • 12+ years in software engineering, cloud operations, or infrastructure roles • 8–10 years of hands-on experience in SRE, Platform Engineering, Cloud Operations, or Production Support for large-scale, distributed SaaS applications • Proven track record leading distributed engineering teams as a player-coach while remaining hands-on with architecture, automation, and incident response • Exceptional troubleshooting skills under pressure and a “Fire Marshal” mindset toward investigation and proactive inspection • Experience working directly with enterprise customers on production reliability, incident communication, and service-level reporting • Must pass customer-mandated background screening for access to critical infrastructure environments • Must be legally authorized to work in the United States • Must complete a drug screen, as applicable • General shift during US business hours and on-call availability for P1/Sev1 incidents • Up to 10% travel to customer sites and team locations • Desired: NERC CIP, SOC2, ISO 27001, or IEC 62443 knowledge/experience • Desired: experience in highly regulated industries such as utilities, financial services, or critical national infrastructure • Desired certifications: AWS DevOps Engineer—Professional or Solutions Architect—Associate/Professional, CKA, SRE Practitioner, and AWS FinOps Practitioner or equivalent

🏖️ Benefits

• Discretionary annual bonus • Medical, dental, vision, and prescription drug coverage • Health Coach from GE Vernova, a 24/7 nurse-based resource • Employee Assistance Program with 24/7 confidential assessment, counseling, and referral services • GE Vernova Retirement Savings Plan • Tax-advantaged 401(k) savings opportunity with company matching contributions and company retirement contributions • Fidelity resources and financial planning consultants • Tuition assistance • Adoption assistance • Paid parental leave • Disability benefits • Life insurance • 12 paid holidays • Permissive time off • Professional development opportunities • Relocation assistance not provided

Apply Now

Similar Jobs

🔥 22 hours ago

Millennium

201 - 500

💼 Consulting

🎖️ Defense

🔒 Cybersecurity

Remote DevSecOps cybersecurity engineer securing DoD software, GitLab CI/CD pipelines, containers, and vulnerability management. Supporting Millennium’s national-security cybersecurity missions through RMF compliance and software assurance.

🔥 23 hours ago

GitLab

1001 - 5000

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

Distinguished Engineer directing GitLab’s CI, CD, Plan, and source-control architecture. Driving scalable AI-native DevOps systems, technical strategy, and engineering quality.

🔥 23 hours ago

Vultr

201 - 500

🤖 Artificial Intelligence

🤝 B2B

🔧 Hardware

Senior SRE maintaining MySQL and PostgreSQL reliability for Vultr’s global cloud infrastructure. Owning monitoring, disaster recovery, incident response, security compliance, and automation.

🇺🇸 United States – Remote

💵 $125k - $135k / year

💰 $329M Debt Financing - Vultr on 2025-06

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 Yesterday

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Site Reliability Engineer supporting Cisco-owned Splunk’s FedRAMP cloud platform. Testing features, automating infrastructure, and leading customer incidents on overnight remote shifts.

🕒 Yesterday

CACI International Inc

10,000+ employees

🎖️ Defense

🏛️ Government

🔒 Cybersecurity

Operating System Deployment Engineer managing secure Windows and Windows Server images for CACI’s DoD enterprise IT services. Automating deployments and supporting physical and virtual infrastructure across 187 bases.

🇺🇸 United States – Remote

💵 $75.2k - $158.1k / year

🔥 Funding within the last year

💰 $500M Post-IPO Debt on 2026-02

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)