Software Engineering Manager – Reliability Engineering, Store Systems

Vaga não está no LinkedIn

🕒 6 dias atrás

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of The Home Depot

The Home Depot

10.000+ funcionários

Fundada em 1978

🏗️ Construção

📦 Logística

🛒 Varejo

💰 Debt Financing em 2007-07

Construction • Logistics • Retail

A Home Depot é um dos principais varejistas de melhoria do lar, oferecendo uma ampla variedade de materiais de construção, produtos para melhoria do lar, artigos para jardim e relacionados. A empresa opera tanto em lojas físicas quanto em uma plataforma online, fornecendo soluções abrangentes para entusiastas de DIY, empreiteiros profissionais e proprietários de imóveis. A Home Depot está comprometida com a diversidade, equidade e inclusão, proporcionando oportunidades de emprego e benefícios a uma força de trabalho diversificada. Além disso, a empresa dá grande ênfase ao atendimento ao cliente e ao engajamento dos colaboradores para manter sua posição como líder confiável no setor de melhoria de residências.

Descrição

• Ensure the resilience, performance, and security of Store Systems and related applications • Lead incident triage, root cause analysis, and blameless postmortems • Drive no-repeat resolutions of systemic problems • Engineer reliability into platforms through automation, change management, incident management, problem management, and destructive testing • Establish and enforce SLOs, SLIs, and SLAs for highly available customer-facing workloads • Develop infrastructure strategy aligned with product goals, dependencies, and end-user requirements • Lead infrastructure configuration, debugging, support, technology roll-outs, and system software/hardware stand-up • Create and optimize specifications for complex technology solutions • Report Systems Engineering progress to leadership • Manage vendor relationships and hardware/software purchase requests • Prioritize escalations and requests from product teams and stakeholders • Produce in-house solution documentation and proactively monitor systems issues • Lead, mentor, coach, recruit, retain, and develop Systems Engineering professionals • Conduct performance reviews and manage individual development plans • Foster collaboration and remove impediments • Advocate for end-user and stakeholder needs • Define and execute the long-term reliability roadmap • Oversee budgets, cloud spending, vendor negotiations, and resource allocation • Lead high-severity incident response and preventative action implementation • Plan capacity, champion chaos engineering, and support disaster recovery planning • Partner with software development, QA, security, and InfoSec teams to integrate reliability and security into the SDLC

🎯 Requisitos

• Must be eighteen years of age or older • Must be legally permitted to work in the United States • Bachelor's degree or equivalent degree in a related field, or equivalent experience • Minimum 5 years of work experience • Technical leadership experience guiding and mentoring SRE, infrastructure, and operations teams • Experience defining and executing reliability roadmaps • Experience recruiting, retaining, developing, and reviewing engineering talent • Experience managing budgets, cloud spending, vendor negotiations, and resource allocation • Expertise defining and enforcing SLOs, SLIs, and SLAs • Experience leading high-severity incidents, blameless postmortems, and preventative actions • Experience with capacity planning, chaos engineering, and disaster recovery • Deep expertise with Google Cloud Platform, AWS, or Microsoft Azure • Advanced knowledge of Terraform, Ansible, Chef, or Puppet • Proficiency with monitoring, logging, and tracing tools such as Datadog, Prometheus, Grafana, Splunk, New Relic, or ELK • Experience overseeing CI/CD pipelines using Jenkins, GitLab CI, or GitHub Actions • Proficiency in one or more of Python, Go, Java, or Bash • Experience partnering with software development, QA, security, and InfoSec teams • Ability to communicate technical metrics and incidents to executive leadership • Experience managing third-party SaaS and infrastructure providers • Knowledge of PCI-DSS, SOC2, HIPAA, and corporate security policies • Knowledge of least-privilege access models and audit logging

🏖️ Benefícios

• Remote/Virtual work arrangement • Overnight travel typically required only 5% to 20% of the time • Mentoring, coaching, and professional development through learning activities and communities of practice • Career development support, including individual development plans and clear career paths • Performance feedback through annual and mid-year reviews

Candidatar-se

Vagas Similares

🕒 6 dias atrás

Sporttrade

11 - 50

💼 Consultoria

📣 Marketing

🎲 Jogos de Azar

Site Reliability Engineer operating Sporttrade’s regulated sports betting exchange across cloud and datacenter infrastructure. Automating operations, improving observability, and leading incident response for a live marketplace.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $170.000 / ano

💰 $36.000.000 Funding Round em 2021-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

NVIDIA

10.000+ funcionários

🏥 Saúde

🏭 Manufatura

🤖 Inteligência Artificial

AI Tools Engineer building LLM and AI/ML systems for NVIDIA’s global GeForce NOW service. Automating incident root-cause analysis and predicting operational trends from production data.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

SitusAMC

5001 - 10000

💼 Consultoria

📦 Logística

🏠 Imobiliário

Site Reliability Engineer operating AWS cloud infrastructure for SitusAMC’s real estate technology solutions. Improving reliability, automation, observability, security, and application migrations.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $95.000 - $135.000 / ano

💰 Private equity em 2020-05

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

GE Vernova

10.000+ funcionários

💼 Consultoria

📦 Logística

🏭 Manufatura

Senior Reliability Engineer improving embedded protection, control, and software products for GE Vernova’s decarbonization mission. Driving testing, field analytics, grid reliability, and cybersecurity compliance.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $162.900 - $244.300 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

SailPoint

1001 - 5000

💼 Consultoria

🏥 Saúde

📦 Logística

Senior Staff DevOps Engineer scaling AWS Kubernetes infrastructure for SailPoint’s identity security platform. Leading enterprise service mesh adoption, PCI-compliant operations, and cloud-native reliability across global teams.

🗣️🇺🇸🇬🇧 Inglês obrigatório