Staff Site Reliability Engineer

🕒 Setembro 14

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $177.000 - $240.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 5%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Fingerprint

Fingerprint

51 - 200 funcionários

Fundada em 2019

🔒 Cibersegurança

🔌 API

☁️ SaaS

💰 $32.000.000 Series B em 2021-11

Cybersecurity • API • SaaS

Fingerprint é uma empresa de tecnologia focada em soluções de detecção e identificação de bots. A plataforma ajuda empresas a garantir um ambiente online seguro ao identificar usuários com precisão e diferenciar entre pessoas reais e bots automatizados. Ela fornece guias para desenvolvedores e configurações de workspace para integração fácil em diversas aplicações, aumentando a segurança e a experiência do usuário.

Descrição

• Serve as Fingerprint's first dedicated Site Reliability Engineer • Work with the Architect on platform design and with Cloud Platform on infrastructure • Partner with every product team on operating their systems • Define SLIs and SLOs for critical request paths and make them visible and actionable • Introduce and coach teams on error budgets • Own reliability metrics used by leadership • Strengthen incident detection, response, communication, postmortems, and follow-up • Improve alert quality, anomaly detection, escalation design, and shared tooling • Lead reliability reviews for high-risk changes and new services • Introduce game days and chaos exercises • Embed with teams on time-boxed reliability engagements • Develop Staff and Lead engineers as reliability leaders • Codify production-readiness, on-call, runbook, and change-safety practices • Partner with the Architect and tech leads to design reliability into systems • Investigate production incidents and write tooling, dashboards, and reference implementations • Lead AI adoption for incident investigation, postmortems, runbooks, observability, and safe AI-assisted operations • Report directly to the VP of Engineering

🎯 Requisitos

• 10+ years of engineering experience • 3+ years as an SRE, production engineer, or reliability-focused Staff engineer operating across multiple teams • Experience owning reliability for a platform • Deep experience with SLI/SLO design and error budgets in practice • Experience driving product-team adoption of reliability practices • Experience leading incident response and postmortems for high-severity, customer-facing incidents • Hands-on knowledge of distributed-systems failure modes, including cache/database saturation, cascading failure, retry storms, capacity limits, degradation, and load shedding • Experience in high-throughput, low-latency environments • Fluency in Kubernetes, AWS, and modern observability tooling such as Datadog or equivalent • Ability to read and write production code in Go, TypeScript, or similar • Experience with infrastructure as code • Track record of leading through influence across teams • Experience coaching engineers to own reliability • Exceptional written communication and documented decision-making • Regular use of AI tools for incident investigation, telemetry analysis, runbooks, postmortems, and tooling • Ability to structure operational data for safe use by humans and AI agents • Pragmatic approach to balancing reliability, delivery, and risk • Must be authorized to work from the home location • Nice to have: experience in fraud detection, identity, payments, or other adversarial real-time domains • Nice to have: multi-region, cell-based, or failure-isolation architecture experience • Nice to have: Elasticsearch, Redis, DynamoDB, or Kafka at scale • Nice to have: familiarity with FinOps and cloud infrastructure reliability/cost trade-offs

🏖️ Benefícios

• Pay transparency • Fully remote work arrangement • Ability to work from almost any country, subject to country restrictions • Inclusive work environment • No visa sponsorship required; teammates may work from their authorized home location

Candidatar-se

Vagas Similares

🕒 Setembro 12

RealTime eClinical Solutions

51 - 200

🏥 Saúde

🧬 Biotecnologia

Principal DevOps Architect owning AWS, Terraform, CI/CD, observability, and AI/ML platforms. Ensuring HIPAA and SOC 2 compliance for clinical research SaaS.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $155.000 - $195.000 / ano

💰 Private Equity Round em 2022-01

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 10

Conga

1001 - 5000

☁️ SaaS

💸 Finanças

🏢 Corporativo

Staff DevOps Engineer building secure cloud infrastructure and CI/CD systems for Conga’s commercial operations software. Managing Kubernetes, observability, IAM, cost optimization, and Terraform automation.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $168.675 - $229.900 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 10

Avanade

10.000+ funcionários

💼 Consultoria

📦 Logística

📣 Marketing

Avanade manager architecting Azure DevOps, GitHub, and AI-enabled software delivery solutions for enterprise clients. Leading DevOps transformation, Copilot adoption, governance, and cloud engineering modernization.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $139.200 - $195.700 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 9

Red Cell Partners

11 - 50

🏥 Saúde

🎖️ Defesa

💼 Consultoria

Staff DevOps Engineer building secure CI/CD and deployment infrastructure for Red Cell’s federal technology companies. Supporting classified environments, ATO activities, and customer deployments.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 8

General Dynamics Information Technology

10.000+ funcionários

💼 Consultoria

🏥 Saúde

📦 Logística

Principal DevOps Engineer building AWS, Kubernetes, data pipelines, and AI platform infrastructure. Supporting GDIT’s technology and mission services for U.S. government, defense, and intelligence agencies.

🗣️🇺🇸🇬🇧 Inglês obrigatório