Staff Software Engineer – Customer Reliability Engineering

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $90k - $190k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of The Home Depot

The Home Depot

10,000+ employees

Founded 1978

🏗️ Construction

📦 Logistics

🛒 Retail

💰 Debt Financing on 2007-07

Construction • Logistics • Retail

The Home Depot is a leading home improvement retailer, offering a wide range of building materials, home improvement products, lawn and garden products, and related services. The company operates both physical stores and an online platform, providing comprehensive solutions for DIY enthusiasts, professional contractors, and homeowners. The Home Depot is committed to diversity, equity, and inclusion, providing employment opportunities and benefits to a diverse workforce. Additionally, the company places a high emphasis on customer service and associate engagement to maintain its position as a trusted leader in the home improvement industry.

📋 Description

• Lead design and implementation of foundational systems for reliability, scalability, performance, and efficiency • Establish enterprise blueprints for observability, automation, and cloud infrastructure architecture • Drive tool selection, configuration, security, resilience, performance tuning, and production monitoring • Define Service Level Objectives, error budgets, and blameless post-incident review practices • Develop, test, deploy, and maintain software • Develop functional and destructive test suites to enable rapid production deployment • Take a broad, global approach to technical issues • Work with Product Teams to ensure user stories are developer-ready, understandable, and testable • Collaborate in agile processes with team members • Field questions from product and engineering teams • Guide junior engineers on modern software development frameworks and lead technical discussions • Identify team gaps and suggest productivity improvements • Drive technical decisions across teams without formal authority • Mentor engineers of all experience levels and elevate engineering standards • Report typically to a Software Engineering Manager or Senior Manager • Typically has 0 direct reports

🎯 Requirements

• Must be eighteen years of age or older • Must be legally permitted to work in the United States • Bachelor's degree program or equivalent degree in a field related to the job • Minimum 3 years of work experience • Preferred 8+ years of relevant professional experience in Cloud Operations, Site Reliability Engineering, DevOps, or Software Engineering in a high-scale, distributed environment • Deep expertise designing, deploying, and operating high-availability, multi-region production architectures on Google Cloud Platform or AWS/Azure • Production-level proficiency in Go, Python, or Java • Deep expertise in observability tools such as Datadog, Prometheus, Grafana, or Splunk • Ability to define and implement SLIs, SLOs, and alerting strategies • Mastery of Infrastructure as Code such as Terraform or CloudFormation • CI/CD pipeline automation experience with GitHub Actions or Jenkins • Advanced operational experience with Kubernetes, container orchestration, and microservices architectures • Experience leading incident response coordination, root-cause analysis, and blameless post-mortems • Ability to lead technical direction of complex, cross-team initiatives • Proven track record mentoring junior and mid-level engineers • Strong communication skills and ability to partner across Product, UX, Architecture, Security, and Engineering teams

🏖️ Benefits

• No travel required • Comfortable indoor working conditions • Frequent opportunity to move about during work

Apply Now

Similar Jobs

🔥 3 hours ago

Merative

1001 - 5000

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Principal SRE directing reliability, observability, and automation for Merative’s medical imaging cloud platforms. Establishing SLOs, resilience, incident management, and infrastructure automation across teams.

🔥 4 hours ago

Bellese Technologies

51 - 200

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

DevSecOps Staff Engineer modernizing CMS Unified Case Management for Bellese, a civic healthcare technology company. Building secure AWS infrastructure, CI/CD pipelines, microservices, and observability solutions.

🕒 Yesterday

SentinelOne

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Staff SRE securing government cloud environments for SentinelOne’s AI-native cybersecurity platform. Leading compliant releases, incident management, observability, and FedRAMP operations.

🕒 Yesterday

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer maintaining reliable AWS and OpenShift production systems for Peraton, a national security and enterprise IT provider. Automating infrastructure, observability, deployments, and incident response.

🕒 Yesterday

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer maintaining AWS, GovCloud, and OpenShift production systems for Peraton’s national security missions. Automating deployments, observability, incident response, and infrastructure resilience.