Search Remote Jobs

Senior Site Reliability Engineer

🔥 1 hour ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of The Access Group

The Access Group

5001 - 10000 employees

💼 Consulting

🏥 Healthcare

🏨 Hospitality

Consulting • Healthcare • Hospitality

The Access Group is a provider of cloud-based, industry-focused business management software and services. It offers modular SaaS solutions — including finance and accounting, HR and payroll, learning and compliance, CRM, ERP, payments and managed IT — tailored to sectors such as charities, education, construction, healthcare, hospitality, recruitment, warehousing and wholesale. The company serves other organisations with integrated, enterprise-grade tools, supported by professional services, customer success and global operations to help customers streamline operations and meet regulatory and sector-specific needs.

📋 Description

• Serve as the senior escalation point for complex P1/P2 production incidents, owning cross-system triage and permanent architectural remediation • Lead platform-level architecture reviews for reliability, scalability, security, and operational standards • Identify systemic failure patterns and translate them into architectural changes, design standards, and platform improvements • Own availability, reliability, performance, and scalability of production systems • Define, track, and improve SLOs, SLIs, and operational KPIs • Develop and maintain Terraform infrastructure-as-code solutions, including modules, state management, and governance • Eliminate operational toil through automation, self-service capabilities, and platform tooling • Build and maintain automation frameworks using Bash, PowerShell, and related scripting technologies • Administer and architect Microsoft Azure solutions, with AWS as a secondary platform • Operate Kubernetes in production, including cluster management and platform maintenance • Manage hybrid-cloud environments, virtual machines, networking, and distributed infrastructure • Maintain Datadog observability and PagerDuty alerting configurations • Design infrastructure controls for PCI-DSS, SOC 1/2, and ISO 27001 compliance and support audits • Partner with Engineering, Product, Security, and Operations on CI/CD, releases, and DevOps maturity • Mentor engineers and influence organizational infrastructure design standards

🎯 Requirements

• 8+ years in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering • Direct ownership of complex production platforms at scale • Senior technical escalation experience for cross-team incidents and architectural remediation • Expert Microsoft Azure infrastructure experience and working knowledge of AWS • Deep Kubernetes production expertise • Advanced Terraform and Infrastructure-as-Code skills, including module design, state management, and governance • Strong Bash scripting and operational automation development • Networking fundamentals including firewalls, DNS, routing, VPN, troubleshooting, and Cloudflare edge services • Active Directory administration and hybrid identity experience • CI/CD pipeline design and deployment workflow improvement experience • Datadog, PagerDuty, or equivalent observability and alerting experience • Experience in compliance-regulated environments, including PCI-DSS, SOC 1/2, and ISO 27001 • Ability to influence across organizational boundaries without direct authority • Applicants must reside within the Eastern or Central time zones • Authorization to work in the U.S. without employer sponsorship is required • Preferred: Puppet or equivalent configuration management administration • Preferred: Microsoft SQL Server administration • Preferred: Meraki firewall policy management • Preferred: AI-driven operational workflows and Model Context Protocol (MCP) development • Preferred: Internal developer platform or platform engineering initiative leadership • Preferred: Large-scale SaaS or high-availability platform support • Preferred: Scala and/or Java infrastructure-level knowledge • Preferred: Azure Solutions Architect Expert, Azure Administrator Associate, AWS Solutions Architect, or CKA certification

🏖️ Benefits

• 22 days paid time off • 11 company paid holidays • Medical insurance • Dental insurance • Vision insurance • 5% 401(k) company match • Range of other selectable benefits • Competitive salary • Development and career progression opportunities • Home-based/remote work arrangement

Apply Now

Similar Jobs

🔥 1 hour ago

AAA Life Insurance Company

501 - 1000

⚕️ Healthcare Insurance

💸 Finance

🛡️ Insurance

Senior DevOps Engineer modernizing AAA Life’s cloud-native infrastructure and middleware. Designing CI/CD, automation, observability, security, and disaster recovery for life insurance systems.

🔥 1 hour ago

US LBM

10,000+ employees

📦 Logistics

🏭 Manufacturing

💼 Consulting

Cybersecurity Engineer securing US LBM’s Azure cloud workloads, DevSecOps pipelines, containers, and data platforms. Embedding security controls, automation, and AI governance across application lifecycles.

🔥 1 hour ago

Jones Lang LaSalle Americas, Inc.

10,000+ employees

🏠 Real Estate

🤝 B2B

💼 Consulting

Cloud Architect & DevOps Manager designing secure Azure/AWS solutions for JLL’s commercial real estate technology. Leading automation, reliability, disaster recovery, and platform engineering teams.

🔥 1 hour ago

Net Health

501 - 1000

🏥 Healthcare

☁️ SaaS

🤖 Artificial Intelligence

DevOps Engineer designing secure AWS platforms and CI/CD automation for Net Health’s healthcare SaaS products. Owning cloud architecture, database performance, observability, security, and cost optimization.

🔥 2 hours ago

Worth AI

11 - 50

💼 Consulting

🛡️ Insurance

🤖 Artificial Intelligence

Senior DevOps Engineer improving Worth AI’s cloud infrastructure, Kubernetes platform, and deployment reliability. Automating infrastructure, strengthening observability, optimizing costs, and enabling engineering teams.