Director, Site Reliability Engineering

Job not on LinkedIn

🔥 2 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Counterpart Health

Counterpart Health

51 - 200 employees

Founded 2024

🏥 Healthcare

🤖 Artificial Intelligence

☁️ SaaS

Healthcare • Artificial Intelligence • SaaS

Counterpart Health is an AI-powered physician enablement platform that delivers data-driven, point-of-care insights to support value-based care. Incubated at Clover Health as Clover Assistant, its proprietary machine-learning models ingest data from 100+ sources, integrate with EHRs, and surface prioritized clinical actions to improve post‑hospitalization follow-up, quality (HEDIS) performance, and cost/risk management for payors, ACOs, and primary care practices. The company offers web-based SaaS integrations, flexible partnership models for value‑based care transformation, and patented ML technologies for diagnosis and medication management.

📋 Description

• Lead and grow our SRE team of ~10 engineers, including hiring, retention, career development, and performance management across multiple time zones (US, HK, NZ). • Build strategic partnerships with product engineering pillars — shifting SRE from reactive, ticket-based support to proactive co-ownership of reliability outcomes. • Scale our multi-tenant infrastructure to support new customer onboarding and growing patient populations. • Own cloud cost management and FinOps practices, building frameworks that balance cost control with reliability and performance. • Champion developer self-service and platform engineering. Build self-service capabilities so product teams can manage routine operations without filing SRE tickets. Establish SLOs/SLIs for critical services and improve alert quality so every page is meaningful. • Ensure the SRE team is fully leveraging AI tooling in their workflows — using tools like Claude Code for IaC generation, log analysis, root cause investigation, and automating repetitive work — at the same level as the rest of engineering.

🎯 Requirements

• You have 6+ years managing an SRE team and 10+ years of hands-on SRE or infrastructure engineering experience. • You're deeply comfortable with our core stack: Kubernetes, GCP (GKE, Cloud SQL, Pub/Sub, GCS), Terraform, Helm, ArgoCD, PostgreSQL, and Prometheus/Grafana. • You have strong programming skills in Python and/or Go, and you're comfortable writing and reviewing infrastructure tooling code — including using AI coding tools to do so. • You have experience with CI/CD pipelines (GitHub Actions) and a track record of building or improving developer tooling and automation. • You have sound build vs. buy judgment — you default to the right answer, not the easiest one, and you're comfortable building internal tooling when existing solutions don't fit. • You have experience leading teams across multiple time zones and a track record of developing engineers into strong technical contributors.

🏖️ Benefits

• Financial Well-Being: Our commitment to attracting and retaining top talent begins with a competitive base salary and equity opportunities. Additionally, we offer a performance-based bonus program, 401k matching, and regular compensation reviews to recognize and reward exceptional contributions. • Physical Well-Being: We prioritize the health and well-being of our employees and their families by providing comprehensive medical, dental, and vision coverage. Your health matters to us, and we invest in ensuring you have access to quality healthcare. • Mental Well-Being: We understand the importance of mental health in fostering productivity and maintaining work-life balance. To support this, we offer initiatives such as No-Meeting Fridays, monthly company holidays, access to mental health resources, and a generous flexible time-off policy. Additionally, we embrace a remote-first culture that supports collaboration and flexibility, allowing our team members to thrive from any location. • Professional Development: Developing internal talent is a priority for Clover. We offer learning programs, mentorship, professional development funding, and regular performance feedback and reviews. • Additional Perks: Employee Stock Purchase Plan (ESPP) offering discounted equity opportunities • Reimbursement for office setup expenses • Monthly cell phone & internet stipend • Remote-first culture, enabling collaboration with global teams • Paid parental leave for all new parents • And much more!

Apply Now

Similar Jobs

🔥 5 hours ago

Mercadona

10,000+ employees

🛒 Retail

🍽️ Food & Beverage

DevOps Prime overseeing GCH’s cloud infrastructure while directing vendor resources and ensuring security. Responsible for CI/CD strategies, observability, and incident response across platforms.

AWS

Cloud

Terraform

🔥 6 hours ago

Finalsite

201 - 500

📚 Education

☁️ SaaS

🤝 B2B

Staff Site Reliability Engineer leading Finalsite's infrastructure evolution and operational excellence practices. Collaborating with engineering leadership to enhance CI/CD and multi-cloud reliability.

🇺🇸 United States – Remote

💵 $180k - $250k / year

💰 Debt financing on 2014-12

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

AWS

Cloud

Google Cloud Platform

Kubernetes

Terraform

🔥 6 hours ago

Whitespace

11 - 50

🎖️ Defense

🏛️ Government

🤖 Artificial Intelligence

Senior DevSecOps Engineer enhancing cybersecurity compliance for federal standards and DoD authorization processes. Leading secure CI/CD implementations and DevSecOps toolchain management for government projects.

Ansible

AWS

Azure

Chef

Cloud

Cyber Security

Docker

Google Cloud Platform

Jenkins

Kubernetes

OpenShift

Python

Terraform

Vault

🔥 7 hours ago

Martian Wall

11 - 50

🎯 Recruiter

💼 Consulting

🤝 B2B

DevOps Architect designing and managing multi-stage CI/CD systems for US-based clients. Strong expertise in cloud and DevOps tool chains is essential.

Ansible

Apache

Chef

Cloud

Docker

Java

Jenkins

Kubernetes

Linux

Perl

Puppet

Python

Ruby

Shell Scripting

🕒 2 days ago

Global Enterprise Services, LLC (GES)

11 - 50

💼 Consulting

📦 Logistics

Reliability Engineer responsible for cloud platform performance and incident response, managing compliance. Requires strong technical expertise and 8 years of experience.

Cloud