Search Remote Jobs

Director of SRE

Job not on LinkedIn

🕒 6 days ago

🇺🇸 United States – Remote

💵 $175k - $200k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Intus Care

Intus Care

11 - 50 employees

💼 Consulting

📣 Marketing

📦 Logistics

💰 $13.1M Venture Round on 2023-01

Consulting • Marketing • Logistics

Intus Care is a healthcare analytics platform that synthesizes healthcare data to identify risks, visualize trends, and optimize care. The company empowers long-term care providers to deliver more effective care to older adults by predicting high-risk patients, reducing expenditures through early risk detection, and improving organizational performance using data-driven insights. Intus Care is particularly beneficial for PACE (Programs of All-Inclusive Care for the Elderly) organizations, providing tools that enable care providers and executives to make informed decisions and proactively manage patient care based on real-time analytics.

📋 Description

• Own and execute the SRE strategy and multi-quarter roadmap across reliability, observability, incident management, QA maturity, and release engineering • Define, measure, and improve SLAs, SLOs, error budgets, uptime, performance, and operational health metrics • Lead production reliability, including monitoring, alerting, on-call operations, incident response, root cause analysis, and MTTR reduction • Establish release readiness standards, deployment safety controls, and quality gates • Manage external SRE vendors and partners, including service delivery, SLA governance, escalations, performance reviews, and compliance expectations • Lead QA engineering strategy focused on automation, regression prevention, test coverage, and reducing escaped production defects • Partner with Security and Engineering leaders on cloud infrastructure, CI/CD pipelines, operational tooling, HIPAA, SOC2, and internal security standards • Oversee Azure AKS, Kubernetes, GitOps workflows, CI/CD pipelines, GitHub Actions, secrets management, access controls, and audit readiness • Drive observability maturity using Grafana, Prometheus, logging platforms, tracing tools, and automated alerting frameworks • Collaborate with Product, Platform, and Engineering teams to embed reliability and quality practices throughout the software development lifecycle • Build, mentor, and scale SRE and QA teams • Drive AI-enabled automation and intelligent tooling to reduce manual toil and improve operational excellence

🎯 Requirements

• 12+ years of SRE, infrastructure, or platform engineering experience • 5+ years of engineering leadership experience • Proven ownership of site reliability for complex, multi-tenant SaaS platforms with demanding availability requirements • Experience defining SLA and SLO frameworks, error budgets, and incident management processes at scale • Experience managing managed infrastructure or SRE service vendors, including SLA governance and performance management • Experience leading QA or quality engineering functions, test automation maturity, and release gate ownership • Strong hands-on experience with Microsoft Azure, preferably including AKS, networking, storage, IAM, and security services • Deep expertise in Kubernetes, containerized workloads, and production-scale distributed systems • Experience with CI/CD pipelines using GitHub Actions, ArgoCD, Terraform, or similar DevOps tooling • Strong background in monitoring, logging, tracing, and observability platforms such as Grafana, Prometheus, Datadog, or Splunk • Experience with scripting and automation using Python, Bash, PowerShell, or similar languages • Strong understanding of release engineering, automated testing frameworks, QA tooling, and shift-left quality practices • Experience supporting SaaS applications with uptime, scalability, and security requirements in regulated industries • Knowledge of HIPAA, SOC2, vulnerability management, access controls, and infrastructure security best practices • Familiarity with databases, APIs, networking, and troubleshooting across modern web application stacks • Strong communication and cross-functional influence skills • Preferred: healthcare technology or regulated SaaS experience; FHIR-native or EMR/EHR architectures; AI-assisted SRE automation; Playwright or equivalent; building internal SRE capability alongside managed services • Must be based in the United States • Position is not eligible for sponsorship

🏖️ Benefits

• Variable compensation component • Stock options • Fully remote, collaborative engineering environment • Direct access to executive leadership • Opportunity to build the SRE function from the ground up • Opportunity to lead a blended team model • Opportunity to work on systems impacting clinical care • AI-assisted software development environment with Claude Code

Apply Now

Similar Jobs

🕒 6 days ago

Veeam Software

1001 - 5000

💼 Consulting

📦 Logistics

☁️ SaaS

Site Reliability Engineer building reliability practices for Veeam’s Government and Sovereign Cloud SaaS platform. Designing Azure infrastructure, observability, automation, and incident-response systems in regulated environments.

🇺🇸 United States – Remote

💵 $138.9k - $231.4k / year

💰 $500M Private Equity Round on 2019-01

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 August 1

Filevine

201 - 500

☁️ SaaS

⚖️ Legal

🤖 Artificial Intelligence

Staff Site Reliability Engineer at Filevine shaping engineering culture and driving reliability practices. Leading technical standards and mentorship within a remote engineering team focused on legal AI technology.

🇺🇸 United States – Remote

💵 $235k - $275k / year

💰 $108M Series D on 2022-04

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 July 31

Aya Healthcare

5001 - 10000

🏥 Healthcare

💼 Consulting

📦 Logistics

Manager of Site Reliability Engineering leading a team for Aya Healthcare's workforce platform. Ensuring product reliability and outstanding user experience through innovative solutions.

🇺🇸 United States – Remote

💵 $230k - $255k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 July 30

TalentWerx

11 - 50

🎯 Recruiter

👥 HR Tech

🤝 B2B

DevOps Engineer IV designing and optimizing deployment solutions for Aether Aerospace. Collaborating with developers to enhance software development processes and ensure system security.

🇺🇸 United States – Remote

💵 $123.6k - $159k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 30

Toast

1001 - 5000

🍽️ Food & Beverage

💼 Consulting

📦 Logistics

Technical leader for release lifecycle and architecting CI/CD framework for Salesforce deployments. Driving automation and technical support in an enterprise ecosystem.