Site Reliability Engineer

Vaga não está no LinkedIn

🕒 Agosto 31

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 11%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Identiq

Identiq

51 - 200 funcionários

Fundada em 2018

💳 Fintech

🛍️ Comércio Eletrônico

🔒 Cibersegurança

💰 $47.000.000 Series A em 2021-03

Fintech • eCommerce • Cybersecurity

A Identiq é uma empresa especializada em otimização de pagamentos e soluções de identificação de clientes. Eles oferecem uma rede privada projetada para aumentar as taxas de aceitação de pagamentos, reduzir fraudes e melhorar as experiências gerais dos clientes sem comprometer dados sensíveis. Sua tecnologia possibilita decisões baseadas em risco por meio do uso de dados próprios, garantindo que informações sensíveis permaneçam seguras e privadas durante todo o processo de validação.

Descrição

• Join the Engineering team as the first dedicated Site Reliability Engineer • Define what reliable means for production systems and establish the SRE practice from zero • Define SLIs and SLOs for core services • Implement Grafana dashboards and burn-rate-based alerting • Establish, own, and continuously improve incident management, including PagerDuty, on-call training, and incident command • Own the observability stack end to end across metrics, logs, traces, RUM, and synthetic checks • Partner with engineering teams to refine SLIs, SLOs, and error budgets and coach teams on SRE and observability best practices • Automate manual and repetitive operational work using infrastructure as code and tooling • Design and run load/performance tests and chaos engineering game days • Drive reliability and infrastructure projects independently at startup speed • Establish documented incident management from detection through blameless postmortems • Reduce alert noise and MTTR and improve confidence in reliability signals for release and investment decisions

🎯 Requisitos

• Bachelor's degree in Computer Science, Computer Engineering, or equivalent formal training, with depth in operating systems, databases, and networking; a rigorous equivalent is accepted • Fundamental systems understanding required for diagnosing novel failures • Daily active use of AI tools to write and debug code, build dashboards and alerts, and increase execution speed • Demonstrated history of independently driving large, ambiguous reliability or infrastructure projects to completion, typically reflecting 5+ years in an SRE, DevOps, or production engineering role • Hands-on implementation of SLI/SLO/error-budget methodology • Strong experience with Grafana and PromQL • Experience with Grafana Alloy for Loki logs and a metrics backend such as Prometheus or Datadog • Experience with OpenTelemetry and a tracing/APM backend such as SigNoz, Uptrace, Tempo, Datadog, or New Relic • Experience with Real User Monitoring (RUM) and synthetic monitoring, such as Grafana Faro, Grafana Synthetic Monitoring, or k6 • Experience designing on-call rotations and incident command practices using PagerDuty or equivalent • Hands-on experience with load/performance frameworks such as Locust, k6, or JMeter • Experience with chaos engineering exercises • Proficiency in Python, Go, or Bash • Hands-on experience with Infrastructure as Code such as Terraform or Ansible • Hands-on experience with Kubernetes • Experience with at least one major cloud platform: AWS, GCP, or Azure • Ability to communicate technical root causes, tradeoffs, implementation details, mitigations, fixes, reliability status, risks, and priorities to technical and business stakeholders • Ability to work independently, resolve ambiguous problems quickly, take ownership, and manage multiple threads under time pressure

🏖️ Benefícios

• Whole-person growth and personal and professional development • Energetic and collaborative environment • Excellent work/life balance • Medical insurance • Dental insurance • Vision insurance • Life insurance • 401k match • Paid time off (PTO) • Two office locations: Downtown Atlanta and Halcyon in Alpharetta

Candidatar-se

Vagas Similares

🕒 Agosto 31

Akamai Technologies

5001 - 10000

🔒 Cibersegurança

Site Reliability Engineer improving reliability, performance, and scalability across Akamai’s distributed cloud and edge platform. Automating operations, strengthening observability, and leading incident response.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $75.700 - $136.300 / ano

💰 Post-IPO Equity em 2001-07

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 31

AXON Networks

201 - 500

💼 Consultoria

📦 Logística

📣 Marketing

Site Reliability Engineer improving reliability across AXON Networks’ AI-driven ISP cloud platform and high-speed routers. Automating NOC operations, observability, incident response and device-management recovery.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $160.000 - $200.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 28

ComPsych

1001 - 5000

🏥 Saúde

💼 Consultoria

📦 Logística

Senior DevSecOps Engineer modernizing cloud infrastructure and CI/CD for ComPsych, a workplace mental-health and absence-management provider. Designing secure automation, observability, and deployment standards across application teams.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 28

Penn Interactive

201 - 500

🎲 Jogos de Azar

🎮 Jogos

🛍️ Comércio Eletrônico

Senior SRE operating Kubernetes and cloud infrastructure for PENN Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production environments.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 28

Summit Racing Equipment

11 - 50

🚘 Automotivo

🛒 Varejo

🛍️ Comércio Eletrônico

DevOps Engineer supporting Summit Racing Equipment’s software delivery and reliability systems. Managing monitoring, CI/CD deployments, SDLC processes, and infrastructure-development team integration.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório