Senior Site Reliability Engineer

🕒 Setembro 18

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $152.000 - $205.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 5%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Fingerprint

Fingerprint

51 - 200 funcionários

Fundada em 2019

🔒 Cibersegurança

🔌 API

☁️ SaaS

💰 $32.000.000 Series B em 2021-11

Cybersecurity • API • SaaS

Fingerprint é uma empresa de tecnologia focada em soluções de detecção e identificação de bots. A plataforma ajuda empresas a garantir um ambiente online seguro ao identificar usuários com precisão e diferenciar entre pessoas reais e bots automatizados. Ela fornece guias para desenvolvedores e configurações de workspace para integração fácil em diversas aplicações, aumentando a segurança e a experiência do usuário.

Descrição

• Own the reliability of core production systems end to end • Define and maintain SLIs and SLOs, dashboards, alerts, and error-budget practices • Improve alert quality and anomaly/correctness detection • Lead incident response, restore service, and write actionable postmortems • Build secure, resilient, and cost-efficient infrastructure with explicit failure-mode handling • Perform load testing, profiling, saturation analysis, and capacity planning • Improve change safety through progressive delivery, automated rollback, pre-production signals, and safe deployment practices • Manage infrastructure through code and configuration, primarily using Terraform • Design, write, and ship software and developer-facing tooling • Run game days and chaos exercises • Partner with product engineering teams on production readiness, capacity, failure modes, rollback plans, runbooks, and on-call handoff • Participate in and improve the on-call rotation • Apply a security lens to engineering work and peer reviews • Serve as the go-to person for difficult production problems and mentor engineers through code review, pairing, and design feedback

🎯 Requisitos

• 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering within primarily cloud-based environments (AWS preferred) • Track record of owning a system end to end • Hands-on experience defining and operating against SLIs, SLOs, and error budgets • Experience leading or serving as a primary responder on high-severity, customer-facing incidents • Depth in distributed-systems failure modes in high-throughput, low-latency environments • Depth in cloud infrastructure fundamentals, including networking, load balancing, containerization (EKS/Kubernetes), and distributed systems • Strong hands-on experience managing infrastructure through code and configuration (Terraform or equivalent) • Solid programming skills in Go, Python, or a comparable language • Fluency with observability tooling such as Datadog, Prometheus, Grafana, or OpenTelemetry • Hands-on experience operating Redis/ElastiCache in production, including cluster/shard management, failover behavior, memory eviction policies, and scaling strategies • Fluency with software engineering best practices, including source control, code review, comprehensive test coverage, and safe deployment • High level of personal ownership and autonomy, with experience working without clearly defined requirements • Pragmatism in balancing reliability and delivery • Strong written and verbal communication in English • AI-native use of AI tools for incident investigation, telemetry analysis, runbooks, and tooling • Must be authorized to work from the home location • Visa sponsorship is not provided

🏖️ Benefícios

• 100% remote work • Ability to join the workforce from almost any country, subject to country restrictions • Visa sponsorship is not provided • Inclusive work environment • CCPA and GDPR notices for applicable residents

Candidatar-se

Vagas Similares

🕒 Setembro 18

Impiricus

11 - 50

🏥 Saúde

💼 Consultoria

🍽️ Alimentos e Bebidas

DevOps engineer building scalable AWS infrastructure and CI/CD for Impiricus’s AI-powered HCP engagement platform. Improving reliability, security, observability, and QA automation.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $100.000 - $130.000 / ano

💰 $3.000.000 Seed Round em 2022-04

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 18

Ondo Finance

51 - 200

₿ Cripto

💳 Fintech

💸 Finanças

Site Reliability Engineer owning reliability, observability, and performance for Ondo Finance’s blockchain-enabled trading platform. Operating Go/Rust services and multi-region AWS Kubernetes infrastructure.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 Initial Coin Offering - Ondo Finance em 2024-01

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 18

Gifthealth

501 - 1000

🏥 Saúde

📦 Logística

💼 Consultoria

DevSecOps Engineer embedding automated security across Gifthealth’s prescription healthcare platform. Building CI/CD, application, infrastructure, container, and Kubernetes security controls.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $115.000 - $165.000 / ano

💰 $40.000.000 Private Equity Round - GiftHealth em 2023-04

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 18

Replit

51 - 200

🤖 Inteligência Artificial

🤝 B2B

Senior Site Reliability Engineer ensuring Replit’s reliable, scalable infrastructure serving millions of developers. Automating operations, observability, incident response, and performance optimization.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 18

TherapyNotes, LLC

51 - 200

💼 Consultoria

⚖️ Jurídico

🏥 Saúde

Site Reliability Engineer improving reliability, observability, and incident response for TherapyNotes’ behavioral health practice-management and EHR SaaS platform. Designing resilient cloud infrastructure for 24×7 production services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $110.000 - $150.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório