Senior Software Engineer – Reliability Engineering

🕒 July 21

🇺🇸 United States – Remote

💵 $90k - $170k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of The Home Depot

The Home Depot

10,000+ employees

Founded 1978

🏗️ Construction

📦 Logistics

🛒 Retail

💰 Debt Financing on 2007-07

Construction • Logistics • Retail

The Home Depot is a leading home improvement retailer, offering a wide range of building materials, home improvement products, lawn and garden products, and related services. The company operates both physical stores and an online platform, providing comprehensive solutions for DIY enthusiasts, professional contractors, and homeowners. The Home Depot is committed to diversity, equity, and inclusion, providing employment opportunities and benefits to a diverse workforce. Additionally, the company places a high emphasis on customer service and associate engagement to maintain its position as a trusted leader in the home improvement industry.

📋 Description

• Develops, tests, deploys, and maintains software, with a clear understanding of the value the software is to provide • Takes on new opportunities and tough challenges with a sense of urgency, high energy and enthusiasm • Consistently achieves results, even under tough circumstances • Develops test suites (functional, destructive, etc) to enable success, rapid deployment of code to production • Takes a broad view when approaching issues • Learns through successful and failed experiment when tackling new problems • Actively seeks ways to grow and be challenged using both formal and informal development channels • Collaborates with other team members in agile processes • Creates new and better ways for the organization to be successful • Works with the Product Team to ensure user stories are valuable, developer ready, easy to understand and testable • Delivers multi-mode communications that convey a clear understanding of the unique needs of different audiences • Adapts approach and demeanor in real time to match the shifting demands of different situations • Relates openly and comfortably with diverse groups of people • Helps grow junior engineers by providing guidance on modern software development frameworks, and leading technical discussions

🎯 Requirements

• 3-5 years of relevant work experience in a related engineering field (Systems, Software, Operational) or a reliability engineering domain • Deep understanding of and extensive experience with ITIL processes and the support/maintenance of production systems, including Change, Incident, and Problem Management • Experience leveraging AI tooling to compress MTTD, MTTM, and MTTR • Extensive experience with common scripting/programming languages (BASH, Python, Golang, Typescript, Java) and data serialization/configuration DSLs (YAML, JSON, HCL) • Extensive experience with infrastructure automation & orchestration tools such as Terraform and Ansible • Extensive experience managing Google Cloud Platform (or equivalent) projects and services, including infrastructure, Compute, Developer Tools, Security, and Identity Access Management • Experience with observability and monitoring tooling such as Prometheus, Grafana, and OpenTelemetry • Strong understanding of container orchestration (Kubernetes/GKE) and modern microservice architectures • Familiarity with both Unix/Linux operating systems • Experience with security tooling (Wiz) & frameworks • Experience designing and executing destructive, performance, and failure-scenario tests, including leading team drills that validate operational readiness • Experience with modern debugging and root cause analysis techniques • Experience with version control systems • Deep understanding of SLOs and core SRE principles and practices • Extensive experience taking a lead role in managing live production incidents and problem management, including reporting business impact to leadership • Strong communication and collaboration skills, with experience producing operational status communications, real-time reporting to diverse stakeholders, documentation, and peer mentorship.

🏖️ Benefits

• Health insurance • 401(k) matching • Flexible work hours

Apply Now

Similar Jobs

🕒 July 20

NextGen Healthcare

1001 - 5000

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Senior Cloud Operations Reliability Engineer driving operational excellence and reliability in cloud services at NextGen Healthcare. Partnering with teams to enhance service availability, resiliency, and automation.

🕒 July 20

Wolters Kluwer

10,000+ employees

🏥 Healthcare

⚖️ Legal

💼 Consulting

Lead DevOps Engineer implementing CI/CD pipelines and cloud infrastructure at Wolters Kluwer. Mentoring a team and improving system reliability through continuous development practices.

🕒 July 20

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Lead reliability workstreams for Akamai's serverless inference platform. Design SRE tooling and automation while mentoring other SREs in cross-team initiatives.

🇺🇸 United States – Remote

💵 $146.4k - $263.6k / year

💰 Post-IPO Equity on 2001-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 July 19

KITC, LLC (8(a) SDB)

11 - 50

🔒 Cybersecurity

🏛️ Government

💼 Consulting

Cloud DevSecOps Engineer III responsible for AWS cloud infrastructure and security automation. Collaborating on federal cybersecurity initiatives in a fully remote role.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 18

Envision Healthcare

10,000+ employees

🏥 Healthcare

👥 B2C

🤝 B2B

DevOps Engineer IV for Envision's Engineering team. Collaborating on implementing advanced CI/CD pipelines and infrastructure as code.