Search Remote Jobs

Senior Reliability Engineer

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $140k - $165k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of hims & hers

hims & hers

201 - 500 employees

Founded 2017

🏥 Healthcare

💼 Consulting

📣 Marketing

Healthcare • Consulting • Marketing

Hims & Hers is an online platform with over 1 million subscribers that connects patients to licensed healthcare professionals across all 50 states in the U. S. The service offers comprehensive support for sexual health, weight loss, hair regrowth, mental health, and skincare through a 100% online process. Clients can receive personalized treatment plans which may include prescription medications, and benefit from free and discreet shipping. Hims & Hers prides itself on providing accessible and affordable healthcare without the need for insurance, offering transparent pricing and support from licensed providers. It aims to empower individuals by making healthcare and treatment conveniently available on their own terms, including mental health support, which includes treatments for anxiety and depression.

📋 Description

• Own reliability for Tier 1 customer journeys, including checkout, telehealth visits, and prescription fulfillment • Define and instrument SLOs, golden signals, and business-level monitors • Improve the ratio of issues detected by monitors versus humans • Lead capacity and resilience work for high-stakes events and steady-state growth • Design load tests and analyze prior events at second-by-second granularity • Investigate database and backend bottlenecks • Harden caching, GraphQL, VPC capacity, and vendor rate limits • Mature FireHydrant, Datadog, and Jira into a connected incident-response pipeline • Automate incident and RCA ticket creation, SLO burn-rate and composite alerting, Tier 1 alert routing, and post-mortem action tracking • Author and maintain runbooks and severity standards • Build and operate AI agents and tooling for OER report generation, RCA drafting, stale action-item detection, monitor and runbook gap detection, and first-pass incident triage • Debug complex cross-boundary issues spanning frontend, API, service mesh, and database layers • Partner with Security and product teams on anomaly detection and response • Produce tooling and metrics for bi-weekly VP-level Operational Excellence reviews and monthly cross-engineering OERs • Drive resulting action items to closure • Document systems, onboard teammates, and coach engineers on SLOs, blameless post-mortems, and on-call practice

🎯 Requirements

• 5+ years as a Software, SRE, Platform, or Infrastructure Engineer • Track record of owning reliability outcomes for production systems that customers depend on • Strong software engineering fundamentals • Ability to solve reliability problems by writing code and building tooling • Comfortable reading application code across the stack • Hands-on depth in observability and SLO engineering, including golden signals, burn-rate alerting, and journey-level monitors • Production experience with AWS, Kubernetes/EKS, Terraform, and PostgreSQL (RDS/Aurora) • Experience running or maturing incident management end to end, including on-call design, escalation policies, incident command, blameless post-mortems, and action-item follow-through • Daily practical use of AI coding and analysis tools such as Claude or Cursor • Judgment about when to trust AI output and when to verify it • Communication skills to explain risk, tradeoffs, and post-incident learnings to engineers and leadership • Preferred: Experience building AI agents or LLM-backed automation for operations • Preferred: Load-testing and performance-engineering experience at meaningful scale, such as k6 • Preferred: Familiarity with service mesh, especially Istio • Preferred: Background in a regulated or healthcare environment • Preferred: Experience designing vendor and partner escalation frameworks with defined severities and response SLAs • Must be legally authorized to work in the U.S. without restriction for any employer • Must not require immigration sponsorship by Hims & Hers to work in the U.S.

🏖️ Benefits

• Competitive salary • Equity compensation • Unlimited PTO • Company holidays • Quarterly mental health days • Medical, dental, and vision benefits • Parental leave • Employee Stock Purchase Program (ESPP) • 401k benefits with employer matching contribution • Offsite team retreats • Claude Enterprise license • Talent-first flexible/remote work approach

Apply Now

Similar Jobs

🔥 34 minutes ago

HopSkipDrive

51 - 200

📦 Logistics

💼 Consulting

✈️ Travel

Senior DevOps manager leading AWS infrastructure, CI/CD, Kubernetes, and AI tooling. Scaling safe student transportation technology for schools and districts nationwide.

🔥 1 hour ago

Penn Mutual

1001 - 5000

Senior DevOps Engineer advancing automation, observability, and reliability for Penn Mutual’s regulated financial services systems. Leading incident response, architecture, CI/CD, and SRE practices.

🔥 3 hours ago

URBN (Urban Outfitters, Anthropologie Group, Free People & Nuuly)

10,000+ employees

👥 B2C

🛒 Retail

👗 Fashion

Senior DevOps Engineer scaling Nuuly's clothing-rental platform infrastructure on GCP, Kubernetes, and Kafka. Automating deployments, monitoring systems, and improving reliability.

🔥 17 hours ago

Ford Motor Company

10,000+ employees

📦 Logistics

💼 Consulting

📣 Marketing

Site Reliability Engineer scaling Ford’s AI-powered observability platform across cloud and on-prem environments. Improving reliability, monitoring, automation, and incident response for connected mobility systems.

🔥 20 hours ago

Entarian

1001 - 5000

🚀 Aerospace

🎖️ Defense

🏛️ Government

Senior Site-Reliability Engineer maintaining reliable Windows production infrastructure for Entarian’s mission-critical engineering solutions. Automating operations, monitoring services, and improving deployment reliability.