Senior Engineering Manager, Site Reliability

🕒 Julho 17

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $195.300 - $270.400 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

👻 Score fantasma 4%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Upstart

Upstart

1001 - 5000 funcionários

🚘 Automotivo

💼 Consultoria

🏥 Saúde

Automotive • Consulting • Healthcare

A Upstart é um dos principais marketplaces de empréstimos com IA, em parceria com bancos e cooperativas de crédito para expandir o acesso a crédito acessível. Ao nos tornarmos uma empresa de capital aberto, estamos agora prontos para alavancar nossa expertise no domínio e revolucionar todos os aspectos da concessão de empréstimos e avaliação de risco de crédito. Recentemente ampliamos nossas ofertas para incluir refinanciamento de automóveis e planejamos atuar em mais verticais à medida que o negócio cresce. Ao utilizar o marketplace de IA da Upstart, bancos e cooperativas de crédito que utilizam a Upstart podem ter taxas de aprovação mais altas e taxas de perda mais baixas, enquanto simultaneamente oferecem a experiência de empréstimo excepcional e digital-first que seus clientes exigem. O marketplace patenteado da Upstart é o primeiro a receber uma carta de não ação do Bureau de Proteção Financeira ao Consumidor relacionada a empréstimos justos. A Upstart não só apoia uma grande força de trabalho remota, mas também possui escritórios em San Mateo, CA; Columbus, OH; e Austin, TX. A maior parte dos que se juntam à Upstart o fazem porque se identificam com nossa missão de permitir acesso a crédito sem esforço baseado no risco real. Se você se sente energizado pelo impacto que pode causar na Upstart, adoraríamos ouvir de você!

Descrição

• Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering • Define a clear charter, priorities, roadmap, and measurable outcomes for the SRE function • Translate strategy into capacity aware plans with explicit trade offs, ownership, milestones, and success measures • Maintain visibility into delivery health, operational risks, and team performance, intervening early when execution drifts • Build a resilient operating model through cross-training, shared context, effective delegation, and clear primary and secondary ownership • Set a high bar for technical quality, operating rigor, and executive communication • Develop engineers and leaders who can independently own complex reliability initiatives • Evolve Upstart’s incident management program to improve detection, response, coordination, communication, and recovery • Establish clear standards for managing high severity incidents and provide visible leadership during critical events • Improve postmortem quality and ensure incident learnings result in durable engineering improvements • Identify recurring failure patterns and drive systemic solutions across teams • Create strong feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction • Improve the quality, accessibility, and trustworthiness of signals used to understand production health • Drive consistent practices across metrics, logs, traces, alerting, and service health • Advance the use of service level objectives and customer impact signals to guide priorities and operational decisions • Reduce detection gaps, noisy alerts, manual investigation, and recurring operational toil • Define measurable reliability outcomes and use data to prioritize investments and communicate impact • Partner with platform and product engineering teams to embed reliability into standard engineering workflows • Establish scalable operational readiness standards for new services, major launches, and architectural changes • Set clear expectations for service ownership, monitoring, capacity, failure handling, and incident response • Identify systemic reliability risks and partner with engineering teams to prioritize and address them • Improve resilience through automation, failure testing, recovery capabilities, and operational safeguards • Build operating mechanisms that turn reviews and analysis into clear decisions, owners, timelines, and sustained follow through • Align stakeholders and dependencies before critical launches and engineering decisions

🎯 Requisitos

• 5+ years of reliability engineering management experience and 7+ years of experience in software engineering, site reliability engineering, infrastructure, or platform engineering • Significant hands-on experience in Site Reliability Engineering, Production Engineering, or an equivalent role responsible for operating and improving production systems • Direct experience managing an SRE, Production Engineering, or equivalent reliability function, including ownership of its strategy, roadmap, operating model, and outcomes • Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations • Experience leading high severity incident response and improving incident management practices at scale • Demonstrated ability to translate strategy into focused, capacity aware plans and deliver measurable outcomes • Track record of hiring, developing, and retaining high performing engineers and engineering leaders • Strong cross-functional leadership and communication, with the ability to turn complex operational data into clear decisions and drive alignment across teams

🏖️ Benefícios

• Competitive compensation, including base pay, bonus opportunities, and annual equity grants that vest quarterly • Retirement benefits to help you plan for the future, including a 401(k) or Group Retirement Savings Plan with a company match of $2 for every $1 contributed, up to $15,000 annually (USD in the US, CAD in Canada) • Employee Stock Purchase Plan (ESPP) with discounted stock purchase options for eligible employees (US only) • Comprehensive health coverage designed to support you and your family, including medical, dental, vision, and wellness resources for US and supplemental health coverage for Canada. • Health Savings Account contributions from Upstart for eligible plans (US only) • Income protection benefits, including life insurance and disability coverage for added financial security • Paid time off, sick leave, and company holidays, in line with local requirements • Paid family and parental leave to support caregiving and major life moments (duration varies by country) • Family-centered benefits to support fertility, parenthood, and caregiving needs • Employee Assistance Program (EAP) offering mental health support and life-centered resources • Financial wellness resources, including access to financial planning tools and a financial concierge service (US Only) • Annual wellness allowance to support your physical and emotional well-being and personal development, based on what matters most to you • Annual productivity allowance to invest in relevant tools and resources you need to do your best work, no matter where you work from • Connection and community through team events, all-company updates, and employee resource groups (ERGs) • Onsite perks, including catered lunches and fully stocked micro-kitchens when working from one of our offices in the Bay Area, Austin, Columbus, and New York City (opening Summer 2026!)

Candidatar-se

Vagas Similares

🕒 Julho 17

Siemens Healthineers

10.000+ funcionários

🏥 Saúde

⚕️ Seguro de Saúde

🧬 Biotecnologia

Network Engineer supporting Varian’s global cloud healthcare infrastructure. Managing Azure, Palo Alto, ExpressRoute, VPN, and hybrid network operations for managed services.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $1.500.000 Grant em 2021-05

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

Azure

Citrix

Cloud

Firewalls

TCP/IP

🕒 Julho 17

Siemens Healthineers

10.000+ funcionários

🏥 Saúde

⚕️ Seguro de Saúde

🧬 Biotecnologia

Network Engineer supporting Siemens Healthineers’ global cloud infrastructure and healthcare applications. Managing Azure networking, Palo Alto firewalls, ExpressRoute, security, and network operations.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $1.500.000 Grant em 2021-05

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

Azure

Citrix

Cloud

Firewalls

TCP/IP

🕒 Julho 16

Miris

11 - 50

☁️ SaaS

🥽 AR/VR

🤝 B2B

Site Reliability Engineer building scalable platforms for 3D/4D content delivery at Miris. Collaborating with teams to ensure system reliability and performance across AR/VR devices.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $102.693 - $287.488 / ano

💰 $26.000.000 Seed Round - MIRIS em 2024-08

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 16

Akamai Technologies

5001 - 10000

🔒 Cibersegurança

Senior Site Reliability Engineer II at Akamai collaborating across software development, operations, and network teams. Responsible for developing standards and tooling for global platform stability.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $146.400 - $263.600 / ano

💰 Post-IPO Equity em 2001-07

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 16

Accelerant

201 - 500

🛡️ Seguros

☁️ SaaS

🤝 B2B

Senior SRE driving reliability and observability across financial data platform at Accelerant. Partnering with engineering to enhance monitoring, metrics, and incident management processes.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $150.000.000 Private Equity Round - Accelerant em 2023-06

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório