Software Engineer II – Reliability Engineering Tooling

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $90k - $150k / year

⏰ Full Time

🟢 Junior

🟡 Mid-level

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of The Home Depot

The Home Depot

10,000+ employees

Founded 1978

🏗️ Construction

📦 Logistics

🛒 Retail

💰 Debt Financing on 2007-07

Construction • Logistics • Retail

The Home Depot is a leading home improvement retailer, offering a wide range of building materials, home improvement products, lawn and garden products, and related services. The company operates both physical stores and an online platform, providing comprehensive solutions for DIY enthusiasts, professional contractors, and homeowners. The Home Depot is committed to diversity, equity, and inclusion, providing employment opportunities and benefits to a diverse workforce. Additionally, the company places a high emphasis on customer service and associate engagement to maintain its position as a trusted leader in the home improvement industry.

📋 Description

• Build and own internal tooling supporting Developers, SREs, and Operations teams. • Own the full lifecycle of applications, including development, testing, deployment, and operations. • Act as an SRE for owned applications, with internal associates as customers. • Automate manual processes and assess site health through meaningful data. • Own SOPs used to deploy and update applications in GCP. • Use AI agents and skills to improve change hygiene and reduce incidents. • Create, deploy, and support production applications. • Collaborate and pair with UX, engineering, product management, and other product team members. • Document, review, and ensure quality and change control standards are met. • Create developer-ready, understandable, and testable user stories with the Product Team. • Write code and scripts to automate infrastructure, monitoring services, test cases, and destructive testing. • Configure and modify programs and commercial off-the-shelf solutions. • Create dashboards, logging, alerting, and proactive responses. • Collaborate in agile processes and improve team effectiveness.

🎯 Requirements

• Must be eighteen years of age or older. • Must be legally permitted to work in the United States. • Bachelor's degree program or equivalent degree in a field of study related to the job. • Minimum 2 years of work experience. • 1-3 years of relevant work experience in a related engineering field or reliability engineering domain. • Working knowledge of ITIL processes and support and maintenance of production systems, including Change, Incident, and Problem Management. • Experience leveraging AI, including prompt engineering and building custom AI agents and skills. • Experience with Golang, JavaScript/TypeScript, and BASH. • Experience writing queries in a relational or noSQL database. • Experience with infrastructure automation, CI/CD processes, Terraform, and GitHub Actions. • Experience working in Google Cloud Platform or equivalent projects and services. • Understanding of Kubernetes/GKE and modern microservice architectures. • Experience with Prometheus, Grafana, and OpenTelemetry. • Exposure to security tooling such as Wiz and security frameworks. • Experience executing destructive, performance, and failure-scenario tests. • Experience with modern debugging and root cause analysis techniques. • Experience with version control systems. • Understanding of SLOs and core SRE principles and practices. • Communication and collaboration skills, including operational status communications, real-time stakeholder reporting, and documentation.

🏖️ Benefits

• Remote/Virtual work arrangement • No travel required

Apply Now

Similar Jobs

🔥 7 hours ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Site Reliability Engineer II operating Akamai's Zero Trust security platform. Automating, monitoring, and optimizing distributed cloud infrastructure for secure, reliable digital experiences.

🔥 10 hours ago

General Dynamics Information Technology

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Cloud Site Reliability Engineer modernizing cloud systems for GDIT’s federal court case-management program. Ensuring resilient operations, monitoring, automation, FinOps, and DevSecOps delivery.

🔥 10 hours ago

Granicus

501 - 1000

🏛️ Government

☁️ SaaS

📋 Compliance

DevOps Engineer II improving CI/CD, cloud infrastructure, and reliability for Granicus government technology platforms. Applying AI and AIOps to deployment, observability, and incident response.

🕒 Yesterday

Gormat

11 - 50

🔒 Cybersecurity

🏛️ Government

🎖️ Defense

Cloud DevOps Engineer developing and integrating cloud-based solutions. Improving system performance, configuration, reliability, and release processes with up to 25% travel.

🕒 Yesterday

Zigabyte

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

DevSecOps Engineer building secure CI/CD pipelines and cloud-native infrastructure. Integrating cybersecurity, automation, and compliance controls for consulting solutions.