Senior Site Reliability Engineer

🔥 46 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Talkiatry

Talkiatry

501 - 1000 employees

Founded 2019

🏥 Healthcare

👥 B2C

Healthcare • B2C

Talkiatry is a virtual psychiatry provider that delivers 100% online psychiatric care and medication management. The platform matches patients to licensed psychiatrists and other clinicians, offers follow-up care with the same provider, and treats conditions such as ADHD, anxiety, depression, bipolar disorder, OCD, PTSD, insomnia and peripartum/postpartum needs. Talkiatry operates in-network with major insurers, provides copay estimates, a patient portal, and resources like quizzes and medication information. The site notes its clinicians average 10 years of experience, represent multiple subspecialties, and speak many languages.

📋 Description

• Define and roll out an SRE practice for a six-team organization: SLOs/SLIs, error budgets, and reliability standards that teams genuinely adopt. • Build and improve observability—metrics, logging, distributed tracing, dashboards, and alerting—so that more incidents are detected by monitoring before anyone outside engineering notices. • Drive down outage frequency by surfacing systemic reliability risks and partnering with teams to remediate them at the root. • Reduce toil through automation, infrastructure-as-code, and self-service tooling that teams can own and extend themselves. • Own the health and usability of our observability tooling, providing documentation and training where necessary. • Run production readiness reviews for new services and partner with engineering leadership on reliability priorities and capacity planning.

🎯 Requirements

• 7+ years in software or infrastructure engineering, with substantial hands-on SRE or production reliability experience. • A track record of reducing incidents and improving detection—the outcomes this role is judged on. • Hands-on experience defining SLOs/SLIs and using error budgets to guide engineering decisions. • Deep observability expertise across metrics, logging, tracing, and alerting (e.g., Datadog, Prometheus, Grafana, or similar). • Strong experience operating production systems on AWS. • Proficiency with infrastructure-as-code (e.g., Terraform) and comfort building automation and tooling (Python, TypeScript, or similar).

🏖️ Benefits

• medical, dental, vision, effective day 1 of employment • 401K with match • generous PTO plus paid holidays • paid parental leave • it all comes back to care: we’re a mental health company, and we put our team’s well-being first • grow your career with us: hone your skills and build new ones with our Learning team as Talkiatry expands

Apply Now

Similar Jobs

🔥 51 minutes ago

Careerswift

2 - 10

👥 HR Tech

🎯 Recruiter

☁️ SaaS

DevOps Engineer responsible for designing, automating, and maintaining CI/CD pipelines. Focus on cloud infrastructure and improving deployment reliability and security.

AWS

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform

🔥 1 hour ago

Multi Media, LLC

51 - 200

💼 Consulting

📣 Marketing

📱 Media

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

Ansible

Cloud

Django

Docker

Flask

Java

Kubernetes

Linux

Laravel

Python

Rust

Switching

Terraform

Go

🔥 2 hours ago

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 United States – Remote

💵 $179.4k - $232.1k / year

💰 $75M Debt Financing - Thumbtack on 2024-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

AWS

Cloud

Distributed Systems

DNS

JavaScript

Linux

Microservices

PHP

Python

SDLC

TCP/IP

Go

🔥 2 hours ago

RTX

10,000+ employees

🚀 Aerospace

🎖️ Defense

🏭 Manufacturing

Senior Principal DevSecOps Engineer designing and implementing DevSecOps platforms for Collins Aerospace. Collaborating with engineers and cybersecurity professionals to enhance software development pipelines and processes.

Docker

Jenkins

Kubernetes

Linux

Maven

Perl

Python

VMware

🔥 2 hours ago

Manulife

10,000+ employees

🛡️ Insurance

💸 Finance

Lead Power Platform Reliability Engineer enhancing enterprise-level solutions through collaboration and mentorship. Shape future data-driven applications and drive cloud integration.

Cloud