Senior SRE

🕒 vor 1 Monat

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

👻 Geisterscore 21%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Accelerant

Accelerant

201 - 500 Mitarbeiter

Gegründet 2018

🛡️ Versicherung

☁️ SaaS

🤝 B2B

💰 €150.000.000 Private Equity Round - Accelerant im 2023-06

Insurance • SaaS • B2B

Accelerant ist eine digitale Risikobörse und Versicherungsplattform, die Managing General Agents (MGAs), Underwriter, Rückversicherer und institutionelles Kapital über einen technologiegestützten, datengesteuerten Marktplatz verbindet. Die Plattform bietet Echtzeitanalysen, Underwriting-Tools, Leistungskennzahlen und operativen Support (aktuarisch, Schadenregulierung, regulatorisch), um die Distribution von Spezialversicherungen zu optimieren und schnellere, transparentere Kapitalbereitstellungen zu ermöglichen. Accelerant positioniert sich als SaaS-Partner für Spezialversicherungsunternehmen mit dem Fokus auf Effizienzsteigerung, Transparenz und profitables Wachstum.

Beschreibung

• Drive the reliability and observability initiative • Own the reliability roadmap end to end. • Prove a repeatable define → emit → ingest → dashboard → alert metric pipeline, set SLOs and error budgets, prioritize the work, and drive execution. • You'll partner with engineering on what we monitor, how, and when — indexing on user impact over low-level infrastructure. • Harden the foundational platform • Take the financial data platform from functional to enterprise-grade, with a focus on availability, performance, and recoverability. • Strengthen deployment paths, straight-through processing, and failover so the monthly close runs faster and cleaner as legacy hops are retired. • Expand observability breadth and depth • Extend instrumentation across the six target systems — Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS (with D365 ledger to follow) — proving both push (OpenTelemetry) and pull (agent) ingestion. • Cover service health (latency, error rates, throughput) and business KPIs (match rate, reconciliation completeness, settlement correctness and latency). • Implement a scalable incident and review process • Build the on-call, alerting, and blameless postmortem process that keeps reliability high as systems and the team grow. • Route alerts Datadog → Incident.io with ServiceNow as the system of record, and set severity standards, escalation norms, and follow-up tracking that actually closes the loop. • Scale automation, auditability, and reduce toil • Build the tooling that automates routine operations, self-heals common failures, and surfaces signal over noise. • Establish data lineage and retention, and validate reliability at scale — 5,000+ transactions before go-live — through auto-remediation, capacity planning, and actionable dashboards. • Build specialized SRE agents using Cursor AI • Design and ship AI agents for incident triage, log analysis, and root-cause investigation (to name a few). • Use Cursor as your build environment. • Treat the agents as products solving specific problems. • Host SRE agents on the AI fabric • Partner with the AI platform team to deploy your agents on the org's AI fabric. Make them discoverable, governed, and reusable across functions.

🎯 Anforderungen

• Proven experience designing, operating, and scaling reliable production systems. • Deep hands-on expertise with modern observability tooling — Datadog, Prometheus/Grafana, and OpenTelemetry — including both push and pull ingestion patterns. • Strong background defining SLIs, SLOs, and error budgets — and translating them into business-level KPIs, not just infrastructure metrics. • Experience operating data platforms (Snowflake, Fabric) and enterprise integration layers (MuleSoft) alongside enterprise SaaS such as D365 (F&O and/or Power Apps). • Hands-on incident management experience with tools like Incident.io and ServiceNow, and a track record of running effective on-call and postmortem practices. • Hands-on experience building with LLMs and AI coding assistants — Cursor in particular. Bonus if you've built and deployed agents. • Ability to define reliability strategy, reliability targets, and operational metrics — and defend them to engineering leadership and the business. • Strong communication skills — you can explain a root cause to a junior engineer and a reliability risk to a product lead. • Demonstrated bias for action and ability to operate autonomously in ambiguous, fast-changing environments.

🏖️ Vorteile

• Competitive salary • Flexible working hours • Professional development budget • Home office setup allowance • Global team events

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Monat

Vytalize Health

201 - 500

🏥 Gesundheitswesen

☁️ SaaS

⚕️ Krankenversicherung

Data Reliability Engineer ensuring operational health of healthcare data pipelines at Vytalize Health. Focused on data quality, compliance, and collaboration across teams.

🇺🇸 Vereinigte Staaten – Remote

💰 €100.000.000 Series C - Vytalize Health im 2023-02

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

11:11 SYSTEMS

201 - 500

💼 Beratung

🏥 Gesundheitswesen

📦 Logistik

Infrastructure Deployment Engineer managing deployment projects across global data centers. Leading cross-functional teams to ensure timely and standardized execution of infrastructure projects.

🇺🇸 Vereinigte Staaten – Remote

💵 $94.000 - $130.500 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Granicus

501 - 1000

🏛️ Regierung

☁️ SaaS

📋 Compliance

DevOps Engineer II automating cloud infrastructure, CI/CD, monitoring, and reliability for Granicus, a government technology solutions provider. Applying AI, MCP tools, and AIOps to improve engineering operations.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Raya

51 - 200

🌍 Soziale Wirkung

👥 B2C

📱 Medien

DevSecOps Engineer improving AWS/EKS security and driving collaboration between DevOps and engineering teams at Raya. Focused on hardening systems and closing security findings across the platform.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

SambaNova Systems

201 - 500

🤖 Künstliche Intelligenz

🔧 Hardware

🏢 Unternehmen

Forward Deployment Engineer embedding with enterprise customers to design and deploy GenAI applications. Collaborating across strategic product offerings to drive value implementation.

🗣️🇺🇸🇬🇧 Englisch erforderlich