Site Reliability Engineering Lead

🔥 12 hours ago

🦌 Connecticut, Florida, +4 more states – Remote

infoinfo

💵 $118.3k - $219.8k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 4%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of RELX

RELX

10,000+ employees

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Consulting • Healthcare • Insurance

RELX is a global provider of information-based analytics and decision tools for professional and business customers. The company focuses on enabling its clients to make better decisions, improve results, and enhance productivity by leveraging advanced technology and data. RELX serves various sectors, including Risk, Scientific, Technical & Medical, Legal, and Exhibitions, by offering specialized information and analytical tools that facilitate critical decision-making. The company is committed to corporate responsibility and delivering societal benefit through its products by contributing to scientific advancement, legal justice, and effective market transactions.

📋 Description

• Manage, mentor, and grow a team of SREs; conduct 1:1s, performance reviews, and career development planning • Own hiring, onboarding, and team capacity/resourcing decisions • Set team goals, prioritize backlog, and drive planning • Foster a blameless post-incident culture and cross-team collaboration with Dev, Security, and Product • Lead reliability initiatives across infrastructure and services • Drive incident response activities and continuous service improvement • Champion automation and operational excellence across the platform • Support the development of scalable, secure, and resilient cloud-native environments • Provide line management to a small to medium-sized team, including performance management, pay, and recruitment authority • Lead post-mortem reviews and ensure timely production of RCAs

🎯 Requirements

• Expert knowledge of Kubernetes, including cluster architecture, upgrades, autoscaling, security hardening, and troubleshooting at scale • Expert experience with Terraform, including modular IaC design, state management, multi-environment provisioning, and policy-as-code • Deep knowledge of Azure Cloud, including compute, networking, identity (AAD), storage, and cost optimization • Experience designing and scaling CI/CD pipelines using GitHub Actions, release strategies, and rollback automation • Experience with observability platforms including Prometheus, Grafana, OpenTelemetry, and SLO/SLA/error-budget management • Strong automation skills focused on eliminating toil through self-healing systems and infrastructure automation • Advanced proficiency in Python, Bash, and/or PowerShell for tooling and automation • Deep understanding of networking concepts including TCP/IP, DNS, load balancing, VPNs, and cloud-native networking • Experience in SRE, DevOps, or Infrastructure roles, including experience leading engineering teams • Proven track record leading incident response and driving reliability improvements

🏖️ Benefits

• Annual incentive bonus • Country-specific benefits • Disability and accommodation support during hiring

Apply Now

Similar Jobs

🔥 13 hours ago

ERPA

501 - 1000

🏥 Healthcare

📦 Logistics

📣 Marketing

DevOps Engineer automating PeopleSoft, testing, and cloud infrastructure for ERPA’s enterprise application managed services. Building Jenkins and Ansible solutions across enterprise environments.

🔥 16 hours ago

Instacart

1001 - 5000

🍽️ Food & Beverage

📦 Logistics

🛍️ eCommerce

Engineering Manager leading Site Reliability Engineers at Instacart, a grocery technology marketplace. Improving platform reliability, scalability, incident response, and operational excellence.

🔥 17 hours ago

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Cisco SRE technical leader migrating specialized Kubernetes workloads from AWS to internal platforms. Operating globally distributed services and improving Kubernetes reliability, networking, and scalability.

🔥 17 hours ago

Databento

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Site Reliability Engineer maintaining reliability, performance, and observability for Databento’s next-generation financial market-data platform. Improving deployments, incident response, and backend infrastructure.

🔥 17 hours ago

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior SRE building Python, AWS, and Terraform reliability solutions for Peraton’s national security missions. Improving cloud platform reliability through automation, observability, and incident management.