Senior Site Reliability Engineer

🔥 1 hour ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Talkiatry

Talkiatry

501 - 1000 employees

Founded 2019

🏥 Healthcare

👥 B2C

Healthcare • B2C

Talkiatry is a virtual psychiatry provider that delivers 100% online psychiatric care and medication management. The platform matches patients to licensed psychiatrists and other clinicians, offers follow-up care with the same provider, and treats conditions such as ADHD, anxiety, depression, bipolar disorder, OCD, PTSD, insomnia and peripartum/postpartum needs. Talkiatry operates in-network with major insurers, provides copay estimates, a patient portal, and resources like quizzes and medication information. The site notes its clinicians average 10 years of experience, represent multiple subspecialties, and speak many languages.

📋 Description

• Define and roll out an SRE practice for a six-team organization: SLOs/SLIs, error budgets, and reliability standards that teams genuinely adopt. • Build and improve observability—metrics, logging, distributed tracing, dashboards, and alerting—so that more incidents are detected by monitoring before anyone outside engineering notices. • Drive down outage frequency by surfacing systemic reliability risks and partnering with teams to remediate them at the root. • Reduce toil through automation, infrastructure-as-code, and self-service tooling that teams can own and extend themselves. • Own the health and usability of our observability tooling, providing documentation and training where necessary. • Run production readiness reviews for new services and partner with engineering leadership on reliability priorities and capacity planning.

🎯 Requirements

• 7+ years in software or infrastructure engineering, with substantial hands-on SRE or production reliability experience. • A track record of reducing incidents and improving detection—the outcomes this role is judged on. • Hands-on experience defining SLOs/SLIs and using error budgets to guide engineering decisions. • Deep observability expertise across metrics, logging, tracing, and alerting (e.g., Datadog, Prometheus, Grafana, or similar). • Strong experience operating production systems on AWS. • Proficiency with infrastructure-as-code (e.g., Terraform) and comfort building automation and tooling (Python, TypeScript, or similar).

🏖️ Benefits

• medical, dental, vision, effective day 1 of employment • 401K with match • generous PTO plus paid holidays • paid parental leave • it all comes back to care: we’re a mental health company, and we put our team’s well-being first • grow your career with us: hone your skills and build new ones with our Learning team as Talkiatry expands

Apply Now

Similar Jobs

🔥 1 hour ago

Careerswift

2 - 10

👥 HR Tech

🎯 Recruiter

☁️ SaaS

DevOps Engineer responsible for designing, automating, and maintaining CI/CD pipelines. Focus on cloud infrastructure and improving deployment reliability and security.

🔥 2 hours ago

Multi Media, LLC

51 - 200

💼 Consulting

📣 Marketing

📱 Media

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

🔥 2 hours ago

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 United States – Remote

💵 $179.4k - $232.1k / year

💰 $75M Debt Financing - Thumbtack on 2024-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 2 hours ago

RTX

10,000+ employees

🚀 Aerospace

🎖️ Defense

🏭 Manufacturing

Senior Principal DevSecOps Engineer designing and implementing DevSecOps platforms for Collins Aerospace. Collaborating with engineers and cybersecurity professionals to enhance software development pipelines and processes.

🔥 3 hours ago

Manulife

10,000+ employees

🛡️ Insurance

💸 Finance

Lead Power Platform Reliability Engineer enhancing enterprise-level solutions through collaboration and mentorship. Shape future data-driven applications and drive cloud integration.