Senior Site Reliability Engineer

🕒 Março 23

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $200.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Backblaze

Backblaze

201 - 500 funcionários

Fundada em 2007

🛍️ Comércio Eletrônico

🏢 Corporativo

💰 $5.000.000 Series A em 2012-07

Cloud Storage • eCommerce • Enterprise

A Backblaze é uma empresa de armazenamento em nuvem que oferece soluções de backup de dados escaláveis e seguras tanto para empresas quanto para indivíduos. Seu serviço B2 Cloud Storage oferece armazenamento de objetos compatível com S3, permitindo que os usuários protejam e gerenciem seus dados com preços transparentes. A Backblaze é especialista em serviços de backup automáticos e ilimitados para sistemas de computador, garantindo opções de proteção e recuperação de dados para os usuários, além de suportar a integração com aplicativos para funcionalidades aprimoradas.

Descrição

• Own and drive the availability, durability, and performance of critical services across all production environments. • Lead and champion complex projects from problem discovery through complete, cross-functional resolution, demonstrating high-level technical ownership. • Define, establish, and enforce service health standards, including working with engineering leadership to implement SLIs, SLOs, and error budget policies for multiple services. • Lead critical incident response and post-incident reviews, translating findings into strategic, long-term service improvements and architectural changes. • Mentor others and act as a subject matter expert in following and evolving established ITIL/OSS processes (incident, change, problem, and capacity management). • Design and architect scalable automation solutions to eliminate toil and improve the efficiency of operational tasks across the entire platform. • Drive the strategic direction of monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, Catchpoint, ELK), and integrate them for comprehensive observability. • Build, maintain, and secure advanced CI/CD pipelines, configuration management, and complex infrastructure as code solutions (Terraform, Ansible, Jenkins). • Write production-grade code (Bash, Python, Go, etc.) to develop new reliability tools and enhance existing systems. • Act as a principal partner to engineering, product, and operations teams, consulting on resilient system design, architecture, and operation. • Lead and formalize the Production Readiness Review (PRR) process, ensuring robust operational handoff for all new services and features. • Lead capacity planning and disaster recovery strategy across critical infrastructure components. • Manage the relationship with vendors and service providers to troubleshoot systemic issues and ensure strict adherence to SLA performance. • Drive the creation of high-quality documentation, proactively share advanced learnings, and cultivate a reliability-first engineering culture across teams. • Own the creation, maintenance, and dissemination of operational playbooks, runbooks, and detailed system documentation. • Proactively identify systemic, recurring issues and architect and drive the implementation of long-term improvements and strategic design action plans. • Be a leading voice in promoting and embedding reliability-focused practices within development and operations teams.

🎯 Requisitos

• Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience). • 8+ years of progressive experience in site reliability, systems engineering, or operations. • Extensive experience designing, scaling, and operating large-scale, production-grade distributed systems. • Expert-level Linux systems administration and advanced troubleshooting skills. • Lead security-minded operations, focusing on system-wide patching, hardening, and proactive vulnerability identification. • Deep mastery of service reliability concepts, including advanced monitoring, complex alerting strategy, leading incident response, and in-depth root cause analysis. • Advanced proficiency in at least one modern scripting/programming language (Python or Go strongly preferred). • Expert knowledge of incident response methodologies and operational best practices. • Proven experience designing and operating container orchestration (Kubernetes, Docker) and microservices concepts required. • Expert experience with Hashicorp products (Nomad, Vault, Terraform) in a production environment.

🏖️ Benefícios

• Healthcare for family, including dental and vision • Competitive compensation and 401K • RSU grants for full-time employees • ESPP program • Flexible vacation policy • Maternity & paternity leave • MacBook Pro to use for work, plus a generous stipend to personalize your workstation • Childcare bonus (human children only) • Fertility treatment and support • Learning & development program • Commuter benefits • Culture that supports a healthy work-life balance

Candidatar-se

Vagas Similares

🕒 Março 21

Arista Networks

1001 - 5000

🏢 Corporativo

📡 Telecomunicações

Site Reliability Engineer at Arista managing CloudVision-as-a-Service platform, ensuring global service reliability, scalability, and stability with a focus on automation and operational excellence.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $101.000 - $161.000 / ano

💰 $2.600.000 Post-IPO Debt em 2015-05

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 21

Chainlink Labs

201 - 500

💸 Finanças

💳 Fintech

🌐 Web 3

Senior Site Reliability Engineer designing infrastructure primitives for decentralized networks. Collaborate on Kubernetes-based control planes and improve operational efficiency.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 20

Latitude.sh

51 - 200

🎮 Jogos

💳 Fintech

Senior Site Reliability Engineer designing and implementing tools for reliable cloud infrastructure. Collaborating with teams to enhance system observability and incident response.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 20

NBCUniversal

10.000+ funcionários

📱 Mídia

Staff Software Engineer overseeing day-to-day operational support of SAP BTP applications at NBCUniversal. Collaborating with onsite teams to enhance engineering strategies and manage production deployments.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Março 19

Upstart

1001 - 5000

🚘 Automotivo

💼 Consultoria

🏥 Saúde

Senior Software Engineer leading technical direction and large initiatives at Upstart. Focusing on building consumer-facing systems and evolving platform architecture.

🗣️🇺🇸🇬🇧 Inglês obrigatório