Senior SRE

🕒 July 16

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 24%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Accelerant

Accelerant

201 - 500 employees

Founded 2018

🛡️ Insurance

☁️ SaaS

🤝 B2B

💰 $150M Private Equity Round - Accelerant on 2023-06

Insurance • SaaS • B2B

Accelerant is a digital-first risk exchange and insurance platform that connects managing general agents (MGAs), underwriters, reinsurers, and institutional capital through a tech-powered, data-driven marketplace. The platform provides real-time analytics, underwriting tools, performance metrics, and operational support (actuarial, claims, regulatory) to streamline specialty insurance distribution and enable faster, more transparent capital deployment. Accelerant positions itself as a SaaS-style partner for specialty insurance firms, focused on improving efficiency, transparency, and profitable growth.

📋 Description

• Drive the reliability and observability initiative • Own the reliability roadmap end to end. • Prove a repeatable define → emit → ingest → dashboard → alert metric pipeline, set SLOs and error budgets, prioritize the work, and drive execution. • You'll partner with engineering on what we monitor, how, and when — indexing on user impact over low-level infrastructure. • Harden the foundational platform • Take the financial data platform from functional to enterprise-grade, with a focus on availability, performance, and recoverability. • Strengthen deployment paths, straight-through processing, and failover so the monthly close runs faster and cleaner as legacy hops are retired. • Expand observability breadth and depth • Extend instrumentation across the six target systems — Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS (with D365 ledger to follow) — proving both push (OpenTelemetry) and pull (agent) ingestion. • Cover service health (latency, error rates, throughput) and business KPIs (match rate, reconciliation completeness, settlement correctness and latency). • Implement a scalable incident and review process • Build the on-call, alerting, and blameless postmortem process that keeps reliability high as systems and the team grow. • Route alerts Datadog → Incident.io with ServiceNow as the system of record, and set severity standards, escalation norms, and follow-up tracking that actually closes the loop. • Scale automation, auditability, and reduce toil • Build the tooling that automates routine operations, self-heals common failures, and surfaces signal over noise. • Establish data lineage and retention, and validate reliability at scale — 5,000+ transactions before go-live — through auto-remediation, capacity planning, and actionable dashboards. • Build specialized SRE agents using Cursor AI • Design and ship AI agents for incident triage, log analysis, and root-cause investigation (to name a few). • Use Cursor as your build environment. • Treat the agents as products solving specific problems. • Host SRE agents on the AI fabric • Partner with the AI platform team to deploy your agents on the org's AI fabric. Make them discoverable, governed, and reusable across functions.

🎯 Requirements

• Proven experience designing, operating, and scaling reliable production systems. • Deep hands-on expertise with modern observability tooling — Datadog, Prometheus/Grafana, and OpenTelemetry — including both push and pull ingestion patterns. • Strong background defining SLIs, SLOs, and error budgets — and translating them into business-level KPIs, not just infrastructure metrics. • Experience operating data platforms (Snowflake, Fabric) and enterprise integration layers (MuleSoft) alongside enterprise SaaS such as D365 (F&O and/or Power Apps). • Hands-on incident management experience with tools like Incident.io and ServiceNow, and a track record of running effective on-call and postmortem practices. • Hands-on experience building with LLMs and AI coding assistants — Cursor in particular. Bonus if you've built and deployed agents. • Ability to define reliability strategy, reliability targets, and operational metrics — and defend them to engineering leadership and the business. • Strong communication skills — you can explain a root cause to a junior engineer and a reliability risk to a product lead. • Demonstrated bias for action and ability to operate autonomously in ambiguous, fast-changing environments.

🏖️ Benefits

• Competitive salary • Flexible working hours • Professional development budget • Home office setup allowance • Global team events

Apply Now

Similar Jobs

🕒 July 15

Vytalize Health

201 - 500

🏥 Healthcare

☁️ SaaS

⚕️ Healthcare Insurance

Data Reliability Engineer ensuring operational health of healthcare data pipelines at Vytalize Health. Focused on data quality, compliance, and collaboration across teams.

🇺🇸 United States – Remote

💰 $100M Series C - Vytalize Health on 2023-02

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

🕒 July 15

11:11 SYSTEMS

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Infrastructure Deployment Engineer managing deployment projects across global data centers. Leading cross-functional teams to ensure timely and standardized execution of infrastructure projects.

🇺🇸 United States – Remote

💵 $94k - $130.5k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 14

RELX

10,000+ employees

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Site Reliability Engineering Lead guiding cloud-native infrastructure, automation, and security for LexisNexis Risk Solutions. Leading SRE teams across the USA and India with Azure, Terraform, and Kubernetes.

🇺🇸 United States – Remote

💵 $118.3k - $219.8k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 14

SambaNova Systems

201 - 500

🤖 Artificial Intelligence

🔧 Hardware

🏢 Enterprise

Forward Deployment Engineer embedding with enterprise customers to design and deploy GenAI applications. Collaborating across strategic product offerings to drive value implementation.

🕒 July 14

11:11 Systems

1001 - 5000

🏢 Enterprise

🔒 Cybersecurity

🤝 B2B

Infrastructure Deployment Engineer leading deployments across global data centers for 11:11 Systems. Responsible for planning, coordinating, and executing hardware infrastructure projects effectively.

🇺🇸 United States – Remote

💵 $94k - $130.5k / year

💰 Private equity on 2021-10

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)