Software Engineering Manager – Reliability Engineering, Store Systems

🔥 2 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of The Home Depot

The Home Depot

10,000+ employees

Founded 1978

🏗️ Construction

📦 Logistics

🛒 Retail

💰 Debt Financing on 2007-07

Construction • Logistics • Retail

The Home Depot is a leading home improvement retailer, offering a wide range of building materials, home improvement products, lawn and garden products, and related services. The company operates both physical stores and an online platform, providing comprehensive solutions for DIY enthusiasts, professional contractors, and homeowners. The Home Depot is committed to diversity, equity, and inclusion, providing employment opportunities and benefits to a diverse workforce. Additionally, the company places a high emphasis on customer service and associate engagement to maintain its position as a trusted leader in the home improvement industry.

📋 Description

• Ensure the resilience, performance, and security of Store Systems and related applications • Lead incident triage, root cause analysis, and blameless postmortems • Drive no-repeat resolutions of systemic problems • Engineer reliability into platforms through automation, change management, incident management, problem management, and destructive testing • Establish and enforce SLOs, SLIs, and SLAs for highly available customer-facing workloads • Develop infrastructure strategy aligned with product goals, dependencies, and end-user requirements • Lead infrastructure configuration, debugging, support, technology roll-outs, and system software/hardware stand-up • Create and optimize specifications for complex technology solutions • Report Systems Engineering progress to leadership • Manage vendor relationships and hardware/software purchase requests • Prioritize escalations and requests from product teams and stakeholders • Produce in-house solution documentation and proactively monitor systems issues • Lead, mentor, coach, recruit, retain, and develop Systems Engineering professionals • Conduct performance reviews and manage individual development plans • Foster collaboration and remove impediments • Advocate for end-user and stakeholder needs • Define and execute the long-term reliability roadmap • Oversee budgets, cloud spending, vendor negotiations, and resource allocation • Lead high-severity incident response and preventative action implementation • Plan capacity, champion chaos engineering, and support disaster recovery planning • Partner with software development, QA, security, and InfoSec teams to integrate reliability and security into the SDLC

🎯 Requirements

• Must be eighteen years of age or older • Must be legally permitted to work in the United States • Bachelor's degree or equivalent degree in a related field, or equivalent experience • Minimum 5 years of work experience • Technical leadership experience guiding and mentoring SRE, infrastructure, and operations teams • Experience defining and executing reliability roadmaps • Experience recruiting, retaining, developing, and reviewing engineering talent • Experience managing budgets, cloud spending, vendor negotiations, and resource allocation • Expertise defining and enforcing SLOs, SLIs, and SLAs • Experience leading high-severity incidents, blameless postmortems, and preventative actions • Experience with capacity planning, chaos engineering, and disaster recovery • Deep expertise with Google Cloud Platform, AWS, or Microsoft Azure • Advanced knowledge of Terraform, Ansible, Chef, or Puppet • Proficiency with monitoring, logging, and tracing tools such as Datadog, Prometheus, Grafana, Splunk, New Relic, or ELK • Experience overseeing CI/CD pipelines using Jenkins, GitLab CI, or GitHub Actions • Proficiency in one or more of Python, Go, Java, or Bash • Experience partnering with software development, QA, security, and InfoSec teams • Ability to communicate technical metrics and incidents to executive leadership • Experience managing third-party SaaS and infrastructure providers • Knowledge of PCI-DSS, SOC2, HIPAA, and corporate security policies • Knowledge of least-privilege access models and audit logging

🏖️ Benefits

• Remote/Virtual work arrangement • Overnight travel typically required only 5% to 20% of the time • Mentoring, coaching, and professional development through learning activities and communities of practice • Career development support, including individual development plans and clear career paths • Performance feedback through annual and mid-year reviews

Apply Now

Similar Jobs

🔥 8 minutes ago

Sporttrade

11 - 50

💼 Consulting

📣 Marketing

🎲 Gambling

Site Reliability Engineer operating Sporttrade’s regulated sports betting exchange across cloud and datacenter infrastructure. Automating operations, improving observability, and leading incident response for a live marketplace.

🇺🇸 United States – Remote

💵 $150k - $170k / year

💰 $36M Funding Round on 2021-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 1 hour ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

AI Tools Engineer building LLM and AI/ML systems for NVIDIA’s global GeForce NOW service. Automating incident root-cause analysis and predicting operational trends from production data.

🔥 2 hours ago

SitusAMC

5001 - 10000

💼 Consulting

📦 Logistics

🏠 Real Estate

Site Reliability Engineer operating AWS cloud infrastructure for SitusAMC’s real estate technology solutions. Improving reliability, automation, observability, security, and application migrations.

🔥 3 hours ago

GE Vernova

10,000+ employees

💼 Consulting

📦 Logistics

🏭 Manufacturing

Senior Reliability Engineer improving embedded protection, control, and software products for GE Vernova’s decarbonization mission. Driving testing, field analytics, grid reliability, and cybersecurity compliance.

🔥 4 hours ago

SailPoint

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Staff DevOps Engineer scaling AWS Kubernetes infrastructure for SailPoint’s identity security platform. Leading enterprise service mesh adoption, PCI-compliant operations, and cloud-native reliability across global teams.