Senior Site Reliability Engineer

🔥 1 hour ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of ServiceTitan

ServiceTitan

1001 - 5000 employees

Founded 2012

💼 Consulting

📦 Logistics

📣 Marketing

💰 $200M Series G on 2021-06

Consulting • Logistics • Marketing

ServiceTitan is a comprehensive software platform designed for the trades industry, providing solutions to enhance productivity and profitability for businesses. It offers a variety of features including dispatching, scheduling, marketing, reporting, and customer experience tools, tailored for trades like plumbing, HVAC, electrical services, and more. ServiceTitan seeks to empower businesses by optimizing operations, improving cash flow, and delivering superior customer experiences through an all-in-one platform. The software includes real-time data analytics, financing options, and mobile capabilities to support the operational needs of contractors and increase their revenue streams. By consolidating multiple business functions into a single platform, ServiceTitan aims to help contractors grow profitably and efficiently.

📋 Description

• Participate in an on-call rotation, using runbooks and playbooks to diagnose and resolve production issues. • Design, build, and maintain observability dashboards and alerting grounded in Service Level Indicators (SLIs) and Service Level Objectives (SLOs). • Operate and improve our Kubernetes-based compute platform. • Work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems. • Investigate and resolve production incidents, including root-cause analysis and follow-up remediation work. • Partner with product engineering teams to review architecture and infrastructure decisions before they ship. • Build and maintain automation that reduces manual, repetitive operational work across the team. • Write and maintain runbooks and documentation to share on-call knowledge across the team. • Help define non-functional requirements — scalability, availability, performance — for new systems as they're designed. • Collaborate across engineering teams to adopt best practices in reliability and observability. • Contribute to CI/CD pipelines and help teams ship changes safely and quickly.

🎯 Requirements

• 8-10+ years of relevant hands-on experience. • Kubernetes (must-have): strong, hands-on understanding of Kubernetes as a system. • SRE principles: practical experience with SLIs, SLOs, and error budgets. • Cloud engineering & networking: solid grounding in AWS or Azure, including networking fundamentals (subnetting, IP addressing). • Observability: deep experience with at least one modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch). • CI/CD: strong understanding of a CI/CD system — GitHub Actions preferred. • Strong programming skills with the ability to build web applications — ideally with solid working knowledge of .NET and ASP.NET. • We're also open to strong Python (Flask, FastAPI) or Java (Spring) backgrounds. • Experience with distributed systems and their common failure modes (retries, timeouts, cascading failures). • Strong production troubleshooting skills — comfortable diagnosing issues under pressure.

🏖️ Benefits

• Flextime, recognition, and support for autonomous work: Flexible time off with ample learning and development opportunities to continue growing your career. • Company-paid medical, dental, and vision (with 100% employer paid options and 90% coverage for dependents) • FSA and HSA, 401k match, and telehealth options including memberships to One Medical. • Parental leave and support, up to $20k in fertility services (i.e. IUI and IVF), surrogacy, and adoption reimbursement. • On demand maternity support through Maven Maternity, free breast milk shipping through Maven Milk, pet insurance, legal advisory services, financial planning tools, and more.

Apply Now

Similar Jobs

🔥 1 hour ago

Filevine

201 - 500

☁️ SaaS

⚖️ Legal

🤖 Artificial Intelligence

Site Reliability Engineer managing AWS infrastructure at Filevine. Ensuring platform reliability and performance with a focus on automation and scalability.

🔥 1 hour ago

Origami Risk

501 - 1000

🏥 Healthcare

🏗️ Construction

📦 Logistics

Site Reliability Engineer responsible for driving improvements in site reliability and scalability. Collaborating cross-functionally to ensure optimum performance across clients' systems.

🔥 1 hour ago

LMI

1001 - 5000

📦 Logistics

🏥 Healthcare

🎖️ Defense

Senior DevSecOps/Platform Engineer designing and maintaining the Navy logistics platform. Building robust CI/CD pipelines and managing cloud infrastructure in AWS GovCloud.

🔥 5 hours ago

Oddball

51 - 200

💼 Consulting

📦 Logistics

🎖️ Defense

DevOps Engineer working on a pivotal Federal program at Oddball to improve daily lives through quality software. Building and maintaining CI/CD pipelines and managing AWS environments.

🔥 7 hours ago

Pluribus Digital

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Lead DevOps Engineer designing and governing enterprise cloud architecture within a federal environment. Collaborating with engineering teams to ensure compliance and alignment with mission objectives.