Senior Reliability Engineer

🕒 Yesterday

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Symbotic

Symbotic

501 - 1000 employees

Founded 2007

📦 Logistics

💼 Consulting

🍽️ Food & Beverage

Logistics • Consulting • Food & Beverage

Symbotic is an industrial automation company that provides an end-to-end, AI-powered warehouse automation platform combining autonomous mobile robots, advanced vision and sensing, and orchestration software to modernize supply chains. Its system integrates hardware and intelligent software to increase throughput, accuracy, and storage density for retail, grocery, CPG, and wholesale customers, solving labor constraints and operational inefficiencies in distribution and fulfillment operations.

📋 Description

• Lead high-impact RCA investigations for complex production incidents spanning software, infrastructure, industrial controls, and production SOP execution • Chair structured, blameless RCA reviews aligned with ITIL Problem Management — fact-based analysis, clear ownership, timely resolution • Serve as the customer-facing technical lead for RCA discussions, updates, and formal deliverables within defined SLA timelines • Analyze logs, telemetry, and incident trends across many issues to find recurring failure patterns — then influence teams across the organization to eliminate them • Present findings, risks, and recommendations to senior internal leadership and customer stakeholders, backed by data you own end to end • Drive Continuous Service Improvement initiatives that measurably reduce repeat incidents and investigation toil through automation, tooling, and reporting • Mentor teammates and elevate RCA quality standards as a senior individual contributor

🎯 Requirements

• Minimum 8 years supporting complex, business-critical production environments, with a career centered on reliability • Minimum 5 years leading technical RCA, post-incident reviews, or ITIL-aligned Problem Management across software, infrastructure, systems, or industrial technology domains • Strong hands-on troubleshooting and data analysis across large-scale distributed systems, on-prem infrastructure, custom software, logs, telemetry, and incident datasets • A track record of regular, ongoing customer interaction — you can tell us who you worked with, at what level, and how often — and of earning trust with executives, technical and non-technical alike • Proven ability to run multiple high-priority investigations in parallel while influencing cross-functional teams, without direct authority, to close actions on time • Bachelor’s degree in a technical field, or equivalent practical experience

🏖️ Benefits

• medical • dental • vision • disability • 401K • PTO

Apply Now

Similar Jobs

🕒 2 days ago

Istari

11 - 50

🚀 Aerospace

☁️ SaaS

🤖 Artificial Intelligence

DevSecOps Engineer securing AWS and Kubernetes infrastructure for a digital engineering company. Collaborating with platform, infrastructure, and security stakeholders to ensure systems are reliable and secure.

AWS

Cloud

Kubernetes

Linux

Terraform

🕒 2 days ago

Ivanti

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer responsible for deploying and managing SaaS environments in AWS and Azure. Collaborating with cross-departmental teams in a dynamic and empowering environment.

Ansible

Apache

AWS

Azure

Cloud

ElasticSearch

Java

Jenkins

Kafka

Linux

MongoDB

NGINX

Postgres

Python

Redis

Splunk

SQL

Go

.NET

🕒 2 days ago

Bitovi

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

Enterprise Systems Engineer V designing and running AI-driven workflow infrastructure for Avalara. Leading resilient and scalable DevOps practices in a remote-first environment.

AWS

Azure

Cloud

Distributed Systems

DNS

Firewalls

Google Cloud Platform

Grafana

Kubernetes

Linux

Prometheus

Splunk

Terraform

🕒 2 days ago

Clover Health

501 - 1000

🏥 Healthcare

🛡️ Insurance

🤖 Artificial Intelligence

Senior Site Reliability Engineer supporting Counterpart Health's technology infrastructure by developing automation tools and troubleshooting issues. Collaborating with cross-functional teams to maintain a scalable infrastructure platform.

AWS

Azure

Cloud

DNS

Docker

Firewalls

Google Cloud Platform

GRPC

Kubernetes

Linux

Prometheus

Python

Shell Scripting

TCP/IP

Go

🕒 3 days ago

Caris Life Sciences

1001 - 5000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior DevOps Engineer managing and operating Kubernetes/AWS EKS infrastructure for Caris Life Sciences. Designing and automating secure, scalable systems in cloud and on-premises environments.

Ansible

AWS

Cloud

Docker

EC2

Flux

Grafana

Kubernetes

Linux

Microservices

MySQL

NGINX

Node.js

Postgres

Prometheus

Python

Terraform