Search Remote Jobs

Senior Site Reliability Engineer

Job not on LinkedIn

🔥 12 hours ago

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Stratus

Stratus

501 - 1000 employees

Founded 2001

🛡️ Insurance

💼 Consulting

📦 Logistics

Insurance • Consulting • Logistics

Stratus is a B2B technology and consulting firm that helps property & casualty insurers and Managing General Agents modernize legacy core systems, migrate to the cloud, and transform insurance data to enable faster product launches, better analytics, and operational efficiency. It offers domain-focused services including Guidewire and Insurity implementation and optimization, cloud migration roadmaps on Azure, AI readiness and execution, data strategy and governance, application maintenance and QA, platform selection support, and on-demand IT talent via embedded delivery pods. Stratus positions itself as a high-touch partner delivering practical, people-first modernization and talent solutions for insurance technology transformation.

📋 Description

• Define and own service level indicators, objectives, and error budgets for customer-facing services • Build measurement pipelines and publish availability and latency against targets • Build and own production observability, including instrumentation standards, dashboards, and actionable alerting • Operate observability across AKS on Azure, Prometheus, Loki, Tempo, Grafana, Istio, and Flux • Establish and run incident response practices, including on-call rotation, paging paths, severity definitions, incident ownership, escalation, and blameless post-mortems • Define rollback and recovery standards and verify them through regular exercises • Lead performance and capacity engineering, including query and index tuning, connection-pool sizing, capacity modeling, and graceful degradation design • Work with the team operating k6 load, stress, spike, and soak testing • Set endpoint latency thresholds tied to SLOs and make test results a delivery gate • Drive production readiness reviews covering instrumentation, alerting, failure modes, resource limits, and rollback • Track post-mortem remediation items through verified production changes • Contribute to business continuity and disaster recovery planning, including backup/restore validation, failover design, and recovery objectives • Coach engineering teams on instrumenting and operating their own services • Write tooling, automation, instrumentation libraries, and production-code fixes • Work across engineering pods, platform and security functions, and customer-facing teams • Report to the Senior Director, Platform Engineering

🎯 Requirements

• 6+ years of professional engineering experience • At least 3+ years in a dedicated SRE or production engineering role at a B2B SaaS company • Demonstrated ownership of SLOs and error budgets in production • Deep, practical observability skills with Prometheus, Loki, Tempo, and Grafana or comparable tools • Hands-on incident command experience at meaningful severity • Experience building on-call and escalation practice from the ground up • Strong database performance skills: query profiling, index design, connection pooling, and diagnosing saturation under load • MongoDB experience strongly preferred; comparable document or relational depth acceptable • Production Kubernetes experience; AKS preferred • Ability to debug pod scheduling, resource limits, networking, and service mesh behavior; Istio preferred • Solid coding ability in at least one of Go, Python, C#, or TypeScript • Willingness to work in a C#/.NET codebase • Experience with load and performance testing tooling such as k6, JMeter, or Gatling • Production experience on Azure or AWS, with understanding of managed-service failure modes • Fluency with AI-assisted engineering tooling and a track record of designing AI-leveraged workflows • Excellent written communication for post-mortems, runbooks, and reliability reports • Judgment and steadiness under pressure • Experience establishing or maturing an SRE practice preferred • Experience with Sentry or comparable application error-monitoring platforms preferred • Experience operating event-driven and real-time systems preferred • Experience operating MongoDB Atlas at production scale preferred • Experience with durable workflow orchestration such as Temporal preferred • Background in multi-region or multi-zone architecture and DR design preferred • Experience with incident.io, PagerDuty, or comparable incident management platforms preferred • Familiarity with DORA metrics and reliability work inside SOC 2 or NIST 800-171 scope preferred • Experience with legacy monolith reliability preferred • Domain interest in MEP, BIM, AEC, or construction technology preferred • Prior experience in a Series B/growth-stage company preferred

🏖️ Benefits

• Comprehensive and competitive health benefits plan • Matching 401k contributions • 20 days annual PTO • Primarily remote work • Occasional annual team onsites

Apply Now

Similar Jobs

🔥 15 hours ago

Fleetio

51 - 200

📦 Logistics

💼 Consulting

🚗 Transport

Senior Site Reliability Engineer scaling Fleetio’s Ruby on Rails fleet-management platform. Improving infrastructure reliability, database performance, observability, and AI-powered operations.

🔥 16 hours ago

Scribe

51 - 200

☁️ SaaS

⚡ Productivity

🏢 Enterprise

Senior DevOps Engineer scaling AWS, Kubernetes, and deployment systems for Scribe’s workflow intelligence platform. Ensuring reliability, observability, and cost-efficient infrastructure.

🔥 17 hours ago

CACI International Inc

10,000+ employees

🎖️ Defense

🏛️ Government

🔒 Cybersecurity

Operating System Deployment Engineer managing secure Windows images and deployment automation for CACI’s DoD enterprise IT services. Supporting physical and virtual systems across 187 bases.

🇺🇸 United States – Remote

💵 $63.3k - $129.7k / year

🔥 Funding within the last year

💰 $500M Post-IPO Debt on 2026-02

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 18 hours ago

Guidehouse

10,000+ employees

🏥 Healthcare

🎖️ Defense

📦 Logistics

DevOps Engineer automating cloud infrastructure, CI/CD pipelines, and container deployments for Guidehouse government applications. Supporting secure, reliable software delivery across development, QA, and operations.

🔥 21 hours ago

Octus

501 - 1000

💼 Consulting

⚖️ Legal

📚 Education

DevOps Engineer managing cloud infrastructure, CI/CD, security, and reliability for Octus, a global credit intelligence and analytics provider. Automating operations across development, data, SRE, security, and IT teams.