Senior Site Reliability Engineer

Vaga não está no LinkedIn

🕒 Agosto 4

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $165.000 - $185.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 24%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of The Access Group

The Access Group

5001 - 10000 funcionários

💼 Consultoria

🏥 Saúde

🏨 Hospitalidade

Consulting • Healthcare • Hospitality

O Access Group é um provedor de software e serviços de gestão empresarial em nuvem, focado em setores industriais. A empresa oferece soluções SaaS modulares — incluindo finanças e contabilidade, RH e folha de pagamento, aprendizagem e conformidade, CRM, ERP, pagamentos e TI gerenciada — adaptadas a setores como instituições de caridade, educação, construção, saúde, hospitalidade, recrutamento, armazenagem e atacado. A companhia atende outras organizações com ferramentas integradas de nível empresarial, suportadas por serviços profissionais, sucesso do cliente e operações globais para ajudar seus clientes a otimizar operações e atender às necessidades regulatórias e específicas de cada setor.

Descrição

• Serve as the senior escalation point for complex P1/P2 production incidents • Own cross-system triage and lead permanent architectural remediation • Lead platform-level architecture reviews for reliability, scalability, security, and operational standards • Identify systemic failure patterns and translate them into architectural changes and lasting platform improvements • Own availability, reliability, performance, and scalability of production systems • Define, track, and improve SLOs, SLIs, and operational KPIs • Develop and maintain Terraform Infrastructure-as-Code solutions, including module design, state management, and governance • Eliminate operational toil through automation, self-service capabilities, and platform tooling • Build and maintain automation frameworks using Bash, PowerShell, and related scripting technologies • Administer and architect Microsoft Azure solutions, with working knowledge of AWS • Operate Kubernetes in production, including cluster management, workload operations, and platform maintenance • Manage hybrid-cloud environments, including virtual machines, networking, and distributed infrastructure • Maintain and improve Datadog observability and PagerDuty alerting configurations • Champion observability standards across metrics, logging, tracing, and alerting • Design infrastructure controls for PCI-DSS, SOC 1/2, and ISO 27001 compliance • Support evidence collection during audits • Partner with Engineering, Product, Security, and Operations teams on CI/CD, release processes, and DevOps maturity • Mentor peers and junior engineers • Influence organizational engineering standards as the internal technical authority on infrastructure design

🎯 Requisitos

• 8+ years in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering • Direct ownership of complex production platforms at scale • Senior technical escalation experience for cross-team incidents requiring architectural-level decision-making and permanent remediation • Expert-level Microsoft Azure cloud infrastructure experience, with working knowledge of AWS • Deep expertise running Kubernetes in production • Advanced Terraform and Infrastructure-as-Code skills, including module design, state management, and governance • Strong Bash scripting and automation development skills • Networking fundamentals including firewalls, DNS, routing, VPN, and network troubleshooting • Cloudflare edge services experience • Active Directory administration and hybrid identity experience • Hands-on CI/CD pipeline design and deployment workflow improvement experience • Datadog, PagerDuty, or equivalent observability and alerting platform experience • Compliance-regulated environment experience, including PCI-DSS, SOC 1/2, and ISO 27001 • Ability to influence across organizational boundaries without direct authority • Applicants must reside within the Eastern or Central time zones • Authorization to work in the U.S. without employer sponsorship • Preferred: Puppet or equivalent configuration management platform administration • Preferred: Microsoft SQL Server environment administration • Preferred: Meraki firewall policy management • Preferred: AI-driven operational workflows and Model Context Protocol (MCP) development • Preferred: Internal developer platform or platform engineering initiative leadership • Preferred: Large-scale SaaS or high-availability platform support • Preferred: Scala and/or Java application ecosystem knowledge at the infrastructure level • Preferred: Azure Solutions Architect Expert, Azure Administrator Associate, AWS Solutions Architect, or CKA certification

🏖️ Benefícios

• 22 days paid time off • 11 company paid holidays • Medical insurance • Dental insurance • Vision insurance • 5% 401(k) company match • Range of other benefits that you can choose from • Blended approach to office working • Development and career progression opportunities

Candidatar-se

Vagas Similares

🕒 Agosto 4

Net Health

501 - 1000

🏥 Saúde

☁️ SaaS

🤖 Inteligência Artificial

DevOps Engineer designing secure AWS platforms and CI/CD automation for Net Health’s healthcare SaaS products. Owning cloud architecture, database performance, observability, security, and cost optimization.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 3

CLARA Analytics

51 - 200

💼 Consultoria

🏥 Saúde

⚖️ Jurídico

DevOps Engineer at CLARA Analytics improving infrastructure-as-code practices in a fully remote environment. Collaborating with cross-functional teams and automating workflows for an AI-powered analytics platform.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 1

JFrog

1001 - 5000

💼 Consultoria

🏥 Saúde

📦 Logística

Senior DevOps Engineer helping JFrog customers build CI/CD platforms with JFrog, Docker, Kubernetes, and cloud technologies. Influencing product roadmaps and training DevOps communities.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 1

Empower

10.000+ funcionários

💸 Finanças

💳 Fintech

👥 B2C

Senior Data Reliability Engineer operating Empower’s AWS-based financial data platform remotely nationwide. Improving production reliability, incident response, observability, SLAs, and disaster recovery.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 1

Scientific Games

10.000+ funcionários

🎮 Jogos

🤝 B2B

Senior DevOps Engineer enhancing infrastructure and deployment reliability for Scientific Games. Collaborating across teams to scale secure, high-performing platforms and improve delivery efficiency.

🗣️🇺🇸🇬🇧 Inglês obrigatório