Lead Site Reliability Engineer

🕒 Setembro 2

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

👻 Score fantasma 10%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Intellum

Intellum

51 - 200 funcionários

Fundada em 2016

💼 Consultoria

📣 Marketing

🏥 Saúde

Consulting • Marketing • Healthcare

A Intellum é uma empresa focada em transformar o aprendizado em crescimento para as empresas, oferecendo uma plataforma abrangente para educar clientes, parceiros e funcionários. Ela oferece uma gama de soluções, incluindo conteúdo ao vivo e sob demanda, programas de certificação e serviços de consultoria para aprimorar iniciativas educacionais. Com ênfase na ciência do aprendizado, a Intellum apoia empresas na criação de experiências educacionais envolventes que impulsionam a retenção e a receita.

Descrição

• Own and drive infrastructure modernization initiatives, including evolution from legacy compute environments to container-orchestrated infrastructure. • Design and maintain infrastructure as code across multiple cloud providers. • Improve CI/CD systems and deployment tooling for efficient, observable, and recoverable releases. • Provide technical leadership through mentorship, architecture guidance, knowledge sharing, and engineering-practice support. • Establish and evolve SLI and SLO practices, monitoring, alerting, and load-testing capabilities. • Lead platform incident response, troubleshooting, root cause analysis, and corrective actions. • Drive visibility into cloud infrastructure costs and incorporate cost considerations into architecture decisions. • Improve developer experience through infrastructure, development environments, deployment workflows, and production feedback loops. • Partner with Security and Engineering on access controls, infrastructure hardening, compliance, and secure infrastructure practices. • Identify operational and infrastructure risks, recommend priorities, and help drive the Systems Engineering technical roadmap. • Mentor engineers and contribute to developing the Systems Engineering team and technical practices. • Perform other duties as assigned.

🎯 Requisitos

• 8+ years of hands-on experience in infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline, including experience building and operating production systems. • Deep hands-on experience designing, operating, and troubleshooting highly available production infrastructure. • Production experience across more than one major cloud provider, with depth in at least one of AWS or Google Cloud and working fluency in the other. • Significant experience with container orchestration and Kubernetes in production environments, including cluster operations, workload configuration, reliability, and troubleshooting. • Experience modernizing production infrastructure, including migrations from VM-based or legacy environments toward containerized or cloud-native architectures. • Strong infrastructure-as-code experience using Terraform or comparable tooling, with an emphasis on repeatability and automation. • Experience building, operating, or significantly improving CI/CD systems and deployment infrastructure. • Strong incident response and troubleshooting capabilities, including experience diagnosing complex distributed-system failures and contributing to effective post-incident review. • Strong Linux administration skills and scripting or programming ability in Ruby, Python, or a comparable language. • Experience working in a SaaS environment where reliability, availability, and production stability are critical. • Ability to collaborate effectively with distributed teams across US and European time zones and participate in an on-call rotation. • Strong communication skills and the ability to provide technical direction, mentor other engineers, and influence infrastructure decisions across teams. • Bachelor's degree in a related field or equivalent practical experience; equivalent experience is genuinely accepted for this role. • Production experience across both AWS and Google Cloud simultaneously is preferred. • Prior leadership or management, cloud cost management or FinOps, Spinnaker/Jenkins, Ruby on Rails, SOC 2, AI-assisted development tooling, or learning technology experience is preferred.

🏖️ Benefícios

• Medical - 100% of employee premiums for selected individual plans • Dental - 100% of employee premiums covered • Vision - 100% of employee premiums covered • LinkedIn Learning • 401(k) plus matching (US Based Only) • Flexible PTO • Calm subscription • Annual Company Retreat • Personal development budgets

Candidatar-se

Vagas Similares

🕒 Setembro 1

The Home Depot

10.000+ funcionários

🏗️ Construção

📦 Logística

🛒 Varejo

Senior software engineer building monitoring, orchestration, and observability automation for The Home Depot’s retail and supply chain operations. Modernizing batch workflows and mentoring automation engineers.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 Debt Financing em 2007-07

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 1

Fuze Health

1001 - 5000

🏥 Saúde

☁️ SaaS

💊 Farmacêutico

Senior DevSecOps Engineer securing AWS/GCP infrastructure, Kubernetes, and CI/CD for Fuze Health’s national pharmacy platform. Driving compliance, resilience, and secure engineering at scale.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $128.000 - $160.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 1

Koniag Government Services

1001 - 5000

🏛️ Governo

🎖️ Defesa

💼 Consultoria

Senior DevSecOps Engineer securing and automating Microsoft Azure platforms. Supporting Koniag’s federal government customers with cloud security, CI/CD, AKS, and reliable digital services.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 31

Identiq

51 - 200

💳 Fintech

🛍️ Comércio Eletrônico

🔒 Cibersegurança

Founding Site Reliability Engineer building observability, incident management, and SRE practices for Incident IQ’s K-12 district workflow platform. Defining SLIs, SLOs, and reliability automation.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $47.000.000 Series A em 2021-03

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 31

Akamai Technologies

5001 - 10000

🔒 Cibersegurança

Site Reliability Engineer improving reliability, performance, and scalability across Akamai’s distributed cloud and edge platform. Automating operations, strengthening observability, and leading incident response.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $75.700 - $136.300 / ano

💰 Post-IPO Equity em 2001-07

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório