Search Remote Jobs

Senior Site Reliability Engineer

Job not on LinkedIn

🕒 August 4

🇺🇸 United States – Remote

💵 $165k - $185k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 24%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of The Access Group

The Access Group

5001 - 10000 employees

💼 Consulting

🏥 Healthcare

🏨 Hospitality

Consulting • Healthcare • Hospitality

The Access Group is a provider of cloud-based, industry-focused business management software and services. It offers modular SaaS solutions — including finance and accounting, HR and payroll, learning and compliance, CRM, ERP, payments and managed IT — tailored to sectors such as charities, education, construction, healthcare, hospitality, recruitment, warehousing and wholesale. The company serves other organisations with integrated, enterprise-grade tools, supported by professional services, customer success and global operations to help customers streamline operations and meet regulatory and sector-specific needs.

📋 Description

• Serve as the senior escalation point for complex P1/P2 production incidents • Own cross-system triage and lead permanent architectural remediation • Lead platform-level architecture reviews for reliability, scalability, security, and operational standards • Identify systemic failure patterns and translate them into architectural changes and lasting platform improvements • Own availability, reliability, performance, and scalability of production systems • Define, track, and improve SLOs, SLIs, and operational KPIs • Develop and maintain Terraform Infrastructure-as-Code solutions, including module design, state management, and governance • Eliminate operational toil through automation, self-service capabilities, and platform tooling • Build and maintain automation frameworks using Bash, PowerShell, and related scripting technologies • Administer and architect Microsoft Azure solutions, with working knowledge of AWS • Operate Kubernetes in production, including cluster management, workload operations, and platform maintenance • Manage hybrid-cloud environments, including virtual machines, networking, and distributed infrastructure • Maintain and improve Datadog observability and PagerDuty alerting configurations • Champion observability standards across metrics, logging, tracing, and alerting • Design infrastructure controls for PCI-DSS, SOC 1/2, and ISO 27001 compliance • Support evidence collection during audits • Partner with Engineering, Product, Security, and Operations teams on CI/CD, release processes, and DevOps maturity • Mentor peers and junior engineers • Influence organizational engineering standards as the internal technical authority on infrastructure design

🎯 Requirements

• 8+ years in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering • Direct ownership of complex production platforms at scale • Senior technical escalation experience for cross-team incidents requiring architectural-level decision-making and permanent remediation • Expert-level Microsoft Azure cloud infrastructure experience, with working knowledge of AWS • Deep expertise running Kubernetes in production • Advanced Terraform and Infrastructure-as-Code skills, including module design, state management, and governance • Strong Bash scripting and automation development skills • Networking fundamentals including firewalls, DNS, routing, VPN, and network troubleshooting • Cloudflare edge services experience • Active Directory administration and hybrid identity experience • Hands-on CI/CD pipeline design and deployment workflow improvement experience • Datadog, PagerDuty, or equivalent observability and alerting platform experience • Compliance-regulated environment experience, including PCI-DSS, SOC 1/2, and ISO 27001 • Ability to influence across organizational boundaries without direct authority • Applicants must reside within the Eastern or Central time zones • Authorization to work in the U.S. without employer sponsorship • Preferred: Puppet or equivalent configuration management platform administration • Preferred: Microsoft SQL Server environment administration • Preferred: Meraki firewall policy management • Preferred: AI-driven operational workflows and Model Context Protocol (MCP) development • Preferred: Internal developer platform or platform engineering initiative leadership • Preferred: Large-scale SaaS or high-availability platform support • Preferred: Scala and/or Java application ecosystem knowledge at the infrastructure level • Preferred: Azure Solutions Architect Expert, Azure Administrator Associate, AWS Solutions Architect, or CKA certification

🏖️ Benefits

• 22 days paid time off • 11 company paid holidays • Medical insurance • Dental insurance • Vision insurance • 5% 401(k) company match • Range of other benefits that you can choose from • Blended approach to office working • Development and career progression opportunities

Apply Now

Similar Jobs

🕒 August 4

Net Health

501 - 1000

🏥 Healthcare

☁️ SaaS

🤖 Artificial Intelligence

DevOps Engineer designing secure AWS platforms and CI/CD automation for Net Health’s healthcare SaaS products. Owning cloud architecture, database performance, observability, security, and cost optimization.

🕒 August 3

CLARA Analytics

51 - 200

💼 Consulting

🏥 Healthcare

⚖️ Legal

DevOps Engineer at CLARA Analytics improving infrastructure-as-code practices in a fully remote environment. Collaborating with cross-functional teams and automating workflows for an AI-powered analytics platform.

🕒 August 1

Empower

10,000+ employees

💸 Finance

💳 Fintech

👥 B2C

Senior Data Reliability Engineer operating Empower’s AWS-based financial data platform remotely nationwide. Improving production reliability, incident response, observability, SLAs, and disaster recovery.

🕒 August 1

Scientific Games

10,000+ employees

🎮 Gaming

🤝 B2B

Senior DevOps Engineer enhancing infrastructure and deployment reliability for Scientific Games. Collaborating across teams to scale secure, high-performing platforms and improve delivery efficiency.

🕒 July 31

Andromeda

11 - 50

🏥 Healthcare

💼 Consulting

🏨 Hospitality

Engineer embedded with teams running large-scale training and inference on GPU clusters in production. Responsible for onboarding, debugging, and improving performance while ensuring reliability.