Senior Site Reliability Engineer

🕒 July 28

🇺🇸 United States – Remote

💵 $137.9k - $221.4k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 17%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of ServiceTitan

ServiceTitan

1001 - 5000 employees

Founded 2012

💼 Consulting

📦 Logistics

📣 Marketing

💰 $200M Series G on 2021-06

Consulting • Logistics • Marketing

ServiceTitan is a comprehensive software platform designed for the trades industry, providing solutions to enhance productivity and profitability for businesses. It offers a variety of features including dispatching, scheduling, marketing, reporting, and customer experience tools, tailored for trades like plumbing, HVAC, electrical services, and more. ServiceTitan seeks to empower businesses by optimizing operations, improving cash flow, and delivering superior customer experiences through an all-in-one platform. The software includes real-time data analytics, financing options, and mobile capabilities to support the operational needs of contractors and increase their revenue streams. By consolidating multiple business functions into a single platform, ServiceTitan aims to help contractors grow profitably and efficiently.

📋 Description

• Participate in an on-call rotation and diagnose and resolve production issues using runbooks and playbooks • Design, build, and maintain observability dashboards and alerting based on SLIs and SLOs • Operate and improve the Kubernetes-based compute platform • Work across Azure/AWS cloud networking and infrastructure • Investigate and resolve production incidents, including root-cause analysis and remediation • Partner with product engineering teams to review architecture and infrastructure decisions • Build and maintain automation to reduce manual operational work • Write and maintain runbooks and documentation • Help define scalability, availability, and performance requirements for new systems • Collaborate across engineering teams to adopt reliability and observability best practices • Contribute to CI/CD pipelines and help teams ship changes safely and quickly

🎯 Requirements

• Strong, hands-on understanding of Kubernetes • Practical experience with SLIs, SLOs, and error budgets • Solid grounding in AWS or Azure • Networking fundamentals, including subnetting and IP addressing • Deep experience with at least one observability stack: OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch • Strong understanding of a CI/CD system; GitHub Actions preferred, with TeamCity, Azure DevOps, or GitLab CI acceptable • Strong programming skills and ability to build web applications • Ideally, working knowledge of .NET and ASP.NET; strong Python with Flask/FastAPI or Java with Spring also accepted • Experience with distributed systems and common failure modes, including retries, timeouts, and cascading failures • Strong production troubleshooting skills • 8–10+ years of relevant hands-on experience • Database experience is nice-to-have and not mandatory

🏖️ Benefits

• Flexible time off • Learning and development opportunities • Comprehensive onboarding program • Leadership training • Programs and events • Bonusly rewards • Peer-nominated awards • Company-paid medical, dental, and vision insurance • 100% employer-paid options and 90% coverage for dependents • FSA and HSA • 401(k) match • Telehealth options, including One Medical memberships • Parental leave and support • Up to $20k in fertility services • Surrogacy and adoption reimbursement • On-demand maternity support through Maven Maternity • Free breast milk shipping through Maven Milk • Pet insurance • Legal advisory services • Financial planning tools • Annual bonus • Equity • Holistic benefits suite • Flexible/autonomous work support

Apply Now

Similar Jobs

🕒 July 28

TensorWave

11 - 50

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

DevOps Software Engineer building and maintaining software layer connecting infrastructure systems for AI workloads. Collaborating with teams to develop automation, orchestration, and communication across platforms.

🕒 July 28

Kings III Emergency Communications

201 - 500

🔐 Security

📡 Telecommunications

🤝 B2B

Senior DevOps/Security Lead for LiftNet managing AWS cloud infrastructure and CI/CD pipelines. Leading information security programs and ensuring SOC 2 compliance for vertical transportation management solutions.

🕒 July 28

Direct Care Innovations

51 - 200

🏥 Healthcare

☁️ SaaS

🏢 Enterprise

Senior DevOps & Cloud Engineer managing Azure infrastructure for Direct Care Innovations' SaaS platform. Collaborating on scalable, reliable, and secure applications and systems while mentoring engineers.

🕒 July 28

Repario

51 - 200

💼 Consulting

📦 Logistics

⚖️ Legal

DevOps Engineer developing code and scripts for automated deployments in legal eDiscovery sector. Collaborating across teams, ensuring security, and mentoring junior staff.

🕒 July 28

IGAMINGHUNT

11 - 50

🎯 Recruiter

🎲 Gambling

🤝 B2B

DevSecOps Engineer responsible for securing cloud infrastructure and CI/CD pipelines in a fintech and iGaming company. Delivering scalable products with a security-first mindset.