Senior SRE

🕒 Julho 16

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Accelerant

Accelerant

201 - 500 funcionários

Fundada em 2018

🛡️ Seguros

☁️ SaaS

🤝 B2B

💰 $150.000.000 Private Equity Round - Accelerant em 2023-06

Insurance • SaaS • B2B

A Accelerant é uma plataforma de troca de risco e seguro digital que conecta agentes gerais gestores (MGAs), subscritores, resseguradores e capital institucional por meio de um mercado orientado por tecnologia e dados. A plataforma fornece análises em tempo real, ferramentas de subscrição, métricas de desempenho e suporte operacional (atuarial, de sinistros, regulatório) para agilizar a distribuição de seguros especializados e permitir uma implantação de capital mais rápida e transparente. A Accelerant se posiciona como uma parceira estilo SaaS para empresas de seguros especializados, concentrando-se em melhorar a eficiência, transparência e crescimento lucrativo.

Descrição

• Drive the reliability and observability initiative • Own the reliability roadmap end to end. • Prove a repeatable define → emit → ingest → dashboard → alert metric pipeline, set SLOs and error budgets, prioritize the work, and drive execution. • You'll partner with engineering on what we monitor, how, and when — indexing on user impact over low-level infrastructure. • Harden the foundational platform • Take the financial data platform from functional to enterprise-grade, with a focus on availability, performance, and recoverability. • Strengthen deployment paths, straight-through processing, and failover so the monthly close runs faster and cleaner as legacy hops are retired. • Expand observability breadth and depth • Extend instrumentation across the six target systems — Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS (with D365 ledger to follow) — proving both push (OpenTelemetry) and pull (agent) ingestion. • Cover service health (latency, error rates, throughput) and business KPIs (match rate, reconciliation completeness, settlement correctness and latency). • Implement a scalable incident and review process • Build the on-call, alerting, and blameless postmortem process that keeps reliability high as systems and the team grow. • Route alerts Datadog → Incident.io with ServiceNow as the system of record, and set severity standards, escalation norms, and follow-up tracking that actually closes the loop. • Scale automation, auditability, and reduce toil • Build the tooling that automates routine operations, self-heals common failures, and surfaces signal over noise. • Establish data lineage and retention, and validate reliability at scale — 5,000+ transactions before go-live — through auto-remediation, capacity planning, and actionable dashboards. • Build specialized SRE agents using Cursor AI • Design and ship AI agents for incident triage, log analysis, and root-cause investigation (to name a few). • Use Cursor as your build environment. • Treat the agents as products solving specific problems. • Host SRE agents on the AI fabric • Partner with the AI platform team to deploy your agents on the org's AI fabric. Make them discoverable, governed, and reusable across functions.

🎯 Requisitos

• Proven experience designing, operating, and scaling reliable production systems. • Deep hands-on expertise with modern observability tooling — Datadog, Prometheus/Grafana, and OpenTelemetry — including both push and pull ingestion patterns. • Strong background defining SLIs, SLOs, and error budgets — and translating them into business-level KPIs, not just infrastructure metrics. • Experience operating data platforms (Snowflake, Fabric) and enterprise integration layers (MuleSoft) alongside enterprise SaaS such as D365 (F&O and/or Power Apps). • Hands-on incident management experience with tools like Incident.io and ServiceNow, and a track record of running effective on-call and postmortem practices. • Hands-on experience building with LLMs and AI coding assistants — Cursor in particular. Bonus if you've built and deployed agents. • Ability to define reliability strategy, reliability targets, and operational metrics — and defend them to engineering leadership and the business. • Strong communication skills — you can explain a root cause to a junior engineer and a reliability risk to a product lead. • Demonstrated bias for action and ability to operate autonomously in ambiguous, fast-changing environments.

🏖️ Benefícios

• Competitive salary • Flexible working hours • Professional development budget • Home office setup allowance • Global team events

Candidatar-se

Vagas Similares

🕒 Julho 15

Empower AI

501 - 1000

🎖️ Defesa

🏥 Saúde

📦 Logística

Sr. DevOps Engineer architecting scalable infrastructure and automating solutions for USCIS in a fully remote role. Mentoring engineers and supporting critical systems with complex challenges.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 15

Encoura

51 - 200

💼 Consultoria

📣 Marketing

📚 Educação

Azure DevOps Engineer responsible for ensuring reliability of large-scale production systems at Encoura. Leading CI/CD initiatives and collaborating with engineering teams on cloud infrastructure.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $116.000 - $128.800 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 15

Vytalize Health

201 - 500

🏥 Saúde

☁️ SaaS

⚕️ Seguro de Saúde

Data Reliability Engineer ensuring operational health of healthcare data pipelines at Vytalize Health. Focused on data quality, compliance, and collaboration across teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $100.000.000 Series C - Vytalize Health em 2023-02

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 15

11:11 SYSTEMS

201 - 500

💼 Consultoria

🏥 Saúde

📦 Logística

Infrastructure Deployment Engineer managing deployment projects across global data centers. Leading cross-functional teams to ensure timely and standardized execution of infrastructure projects.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $94.000 - $130.500 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 15

Ad Hoc LLC

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

DevOps Engineer IV leading CI/CD pipeline development for a tech company focused on public-sector digital services. Mentoring engineers and improving DevOps processes and practices.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $120.000 - $150.000 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório