Senior Site Reliability Engineer

🕒 3 dias atrás

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of ServiceTitan

ServiceTitan

1001 - 5000 funcionários

Fundada em 2012

💼 Consultoria

📦 Logística

📣 Marketing

💰 $200.000.000 Series G em 2021-06

Consulting • Logistics • Marketing

ServiceTitan é uma plataforma de software abrangente projetada para a indústria de serviços, oferecendo soluções para melhorar a produtividade e a rentabilidade das empresas. Apresenta uma variedade de recursos, incluindo despacho, agendamento, marketing, relatórios e ferramentas de experiência do cliente, adaptadas para áreas como encanamento, HVAC, serviços elétricos e mais. A ServiceTitan busca capacitar as empresas ao otimizar operações, melhorar o fluxo de caixa e oferecer experiências superiores aos clientes através de uma plataforma tudo-em-um. O software inclui análises de dados em tempo real, opções de financiamento e capacidades móveis para apoiar as necessidades operacionais dos empreiteiros e aumentar suas fontes de receita. Ao consolidar múltiplas funções de negócios em uma única plataforma, a ServiceTitan visa ajudar os empreiteiros a crescer de forma rentável e eficiente.

Descrição

• Participate in an on-call rotation, using runbooks and playbooks to diagnose and resolve production issues. • Design, build, and maintain observability dashboards and alerting grounded in Service Level Indicators (SLIs) and Service Level Objectives (SLOs). • Operate and improve our Kubernetes-based compute platform. • Work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems. • Investigate and resolve production incidents, including root-cause analysis and follow-up remediation work. • Partner with product engineering teams to review architecture and infrastructure decisions before they ship. • Build and maintain automation that reduces manual, repetitive operational work across the team. • Write and maintain runbooks and documentation to share on-call knowledge across the team. • Help define non-functional requirements — scalability, availability, performance — for new systems as they're designed. • Collaborate across engineering teams to adopt best practices in reliability and observability. • Contribute to CI/CD pipelines and help teams ship changes safely and quickly.

🎯 Requisitos

• 8-10+ years of relevant hands-on experience. • Kubernetes (must-have): strong, hands-on understanding of Kubernetes as a system. • SRE principles: practical experience with SLIs, SLOs, and error budgets. • Cloud engineering & networking: solid grounding in AWS or Azure, including networking fundamentals (subnetting, IP addressing). • Observability: deep experience with at least one modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch). • CI/CD: strong understanding of a CI/CD system — GitHub Actions preferred. • Strong programming skills with the ability to build web applications — ideally with solid working knowledge of .NET and ASP.NET. • We're also open to strong Python (Flask, FastAPI) or Java (Spring) backgrounds. • Experience with distributed systems and their common failure modes (retries, timeouts, cascading failures). • Strong production troubleshooting skills — comfortable diagnosing issues under pressure.

🏖️ Benefícios

• Flextime, recognition, and support for autonomous work: Flexible time off with ample learning and development opportunities to continue growing your career. • Company-paid medical, dental, and vision (with 100% employer paid options and 90% coverage for dependents) • FSA and HSA, 401k match, and telehealth options including memberships to One Medical. • Parental leave and support, up to $20k in fertility services (i.e. IUI and IVF), surrogacy, and adoption reimbursement. • On demand maternity support through Maven Maternity, free breast milk shipping through Maven Milk, pet insurance, legal advisory services, financial planning tools, and more.

Candidatar-se

Vagas Similares

🕒 3 dias atrás

Filevine

201 - 500

☁️ SaaS

⚖️ Jurídico

🤖 Inteligência Artificial

Site Reliability Engineer managing AWS infrastructure at Filevine. Ensuring platform reliability and performance with a focus on automation and scalability.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $175.000 - $195.000 / ano

💰 $108.000.000 Series D em 2022-04

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Origami Risk

501 - 1000

🏥 Saúde

🏗️ Construção

📦 Logística

Site Reliability Engineer responsible for driving improvements in site reliability and scalability. Collaborating cross-functionally to ensure optimum performance across clients' systems.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $100.000 - $120.000 / ano

💰 Private Equity Round em 2018-03

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

LMI

1001 - 5000

📦 Logística

🏥 Saúde

🎖️ Defesa

Senior DevSecOps/Platform Engineer designing and maintaining the Navy logistics platform. Building robust CI/CD pipelines and managing cloud infrastructure in AWS GovCloud.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Oddball

51 - 200

💼 Consultoria

📦 Logística

🎖️ Defesa

DevOps Engineer working on a pivotal Federal program at Oddball to improve daily lives through quality software. Building and maintaining CI/CD pipelines and managing AWS environments.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $125.000 - $160.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Pluribus Digital

51 - 200

💼 Consultoria

🏥 Saúde

📦 Logística

Lead DevOps Engineer designing and governing enterprise cloud architecture within a federal environment. Collaborating with engineering teams to ensure compliance and alignment with mission objectives.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $160.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório