Search Remote Jobs

Senior Site Reliability Engineer II

🔥 0 minutes ago

🌲 North Carolina, Ohio, +2 more states – Remote

infoinfo

💵 $104.9k - $174.7k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 3%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of RELX

RELX

10,000+ employees

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Consulting • Healthcare • Insurance

RELX is a global provider of information-based analytics and decision tools for professional and business customers. The company focuses on enabling its clients to make better decisions, improve results, and enhance productivity by leveraging advanced technology and data. RELX serves various sectors, including Risk, Scientific, Technical & Medical, Legal, and Exhibitions, by offering specialized information and analytical tools that facilitate critical decision-making. The company is committed to corporate responsibility and delivering societal benefit through its products by contributing to scientific advancement, legal justice, and effective market transactions.

📋 Description

• Drive day-to-day operations for the Life Sciences SRE team, supporting priorities set by the SRE Manager • Deliver reliability and toil-reduction initiatives with SREs and contractors • Partner with the Senior Software Architect for Life Sciences to modernize processes, tooling, and development strategy • Collaborate with SREs across Reed Tech to converge on common tools, standards, and practices • Influence service level objectives across Life Sciences production systems • Participate in major incident resolution as an escalation point and guide restoration efforts • Troubleshoot and resolve complex systems and application issues across development and production environments • Improve the SRE framework and contribute to shared SRE knowledge documentation • Create disaster recovery plans • Mentor and prepare junior SREs for on-call readiness • Champion AI-assisted development and operational tooling • Implement reusable dashboards, observability standards, SLOs, and error budgets • Conduct incident analysis, performance optimization, failover testing, production recovery, and blameless post-mortems • Support platform modernization, CI/CD and SDLC improvements, standard platform adoption, and operational health reporting • Promote SRE best practices, coach engineers, identify skill gaps, and lead automation and toil-reduction initiatives

🎯 Requirements

• Experienced SRE who uses knowledge of internal and external issues to improve services and processes • Experience owning complex reliability and toil-reduction projects • Experience acting as an escalation point during major incidents • Microsoft Azure experience • Experience with VMs/VMSS, App Service, Azure Functions, and growing use of Azure Kubernetes Service (AKS) • Azure DevOps Pipelines experience • Source control experience with GitHub; GitHub Actions under evaluation • Azure Monitor and Log Analytics experience; Datadog experience supplemental • Terraform experience • Experience using AI-assisted coding tools such as Claude, GitHub Copilot, or Codex • Interest or hands-on experience applying AI/ML to observability, automation, or incident response • Comfortable experimenting with and evaluating emerging AI tooling • Strong understanding of full-stack observability, incident analysis, and performance optimization • Experience with on-call readiness, mentoring, major incident response, blameless post-mortems, and root-cause analysis • Advanced knowledge of high-availability systems, resiliency patterns, deployment strategies, and recovery practices • Experience with failover testing, production recovery, and automating recovery processes using Infrastructure-as-Code and configuration management tools • Experience supporting platform modernization, CI/CD and SDLC improvements, standard platform capabilities, and operational health reporting • Ability to act as a trusted technical advisor, contribute to modernization initiatives, drive alignment on standards and tooling, and mentor team members • Ability to promote SRE best practices, coach engineers, identify skill gaps, and champion automation and toil-reduction initiatives

🏖️ Benefits

• Comprehensive, multi-carrier program for medical, dental and vision benefits • 401(k) with match • Employee Share Purchase Plan • Wellness platform with incentives • Headspace app subscription • Employee Assistance and Time-off Programs • Short-and-Long Term Disability, Life and Accidental Death Insurance, Critical Illness, and Hospital Indemnity • Family Benefits, including bonding and family care leaves, adoption and surrogacy benefits • Health Savings, Health Care, Dependent Care and Commuter Spending Accounts • Up to two days of paid leave each to participate in Employee Resource Groups and to volunteer with your charity of choice • Flexible hours • Shared parental leave • Study assistance • Sabbaticals

Apply Now

Similar Jobs

🔥 3 hours ago

Avanade

10,000+ employees

💼 Consulting

📦 Logistics

📣 Marketing

Avanade manager architecting Azure DevOps, GitHub, and AI-enabled software delivery solutions for enterprise clients. Leading DevOps transformation, Copilot adoption, governance, and cloud engineering modernization.

🔥 15 hours ago

Akkadian Labs

51 - 200

☁️ SaaS

🏢 Enterprise

📡 Telecommunications

Senior DevOps Engineer scaling secure AWS, hybrid-cloud, and AI infrastructure for Akkadian Labs’ enterprise collaboration automation platform. Improving deployments, observability, reliability, and operational governance.

🔥 17 hours ago

Casper Studios

2 - 10

🤖 Artificial Intelligence

🏢 Enterprise

📱 Media

AI Security & DevOps Engineer securing AI systems from prototype to production for Casper Studios, an AI services firm. Owning security standards, platform controls, and enterprise client approvals.

🔥 20 hours ago

XBOX

10,000+ employees

🎮 Gaming

🔧 Hardware

Senior SRE improving Blizzard’s large-scale data, analytics, and ML platforms. Building automation, Kubernetes infrastructure, and reliability tooling for gaming services.

🔥 20 hours ago

Trumid

51 - 200

💳 Fintech

💸 Finance

☁️ SaaS

Senior DBRE safeguarding Postgres durability, recovery, and performance for Trumid’s fixed-income trading platform. Owning failover, observability, backups, and data resilience.