Site Reliability Engineer

🔥 14 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cracker Barrel

Cracker Barrel

10,000+ employees

Founded 1969

🏨 Hospitality

👥 B2C

🍽️ Food & Beverage

💰 $300M Post-IPO Debt - Cracker Barrel on 2025-06

Hospitality • B2C • Food & Beverage

Cracker Barrel is a vibrant business grounded in a clear belief - the goodness of country hospitality. They are dedicated to serving not just meals, but also creating moments of appreciation and togetherness for their customers. With a focus on country-style hospitality, Cracker Barrel aims to foster a welcoming atmosphere where guests feel valued and connected.

📋 Description

• Design, implement, and maintain reliability practices that improve availability, performance, scalability, resiliency, and operational maturity across digital platforms and supporting services. • Monitor production systems using observability tools, dashboards, logs, metrics, traces, alerts, and synthetic monitoring to identify issues before they impact customers or associates. • Provide Tier-2 and Tier-3 production support for customer-facing and internal digital applications, including web, mobile, commerce, CMS, APIs, integrations, and cloud-hosted services. • Lead and participate in incident response, root-cause analysis, problem management, post-incident reviews, and follow-up actions that reduce recurrence and improve service reliability. • Develop automation, scripts, runbooks, self-healing processes, and operational tools that reduce manual effort, accelerate recovery, and improve consistency across environments. • Partner with development teams to improve CI/CD pipelines, deployment readiness, release validation, rollback procedures, feature monitoring, and environment stability. • Collaborate with infrastructure, cloud, security, architecture, QA, and vendor teams to ensure systems meet company standards for security, privacy, compliance, resiliency, and operational support. • Define and track service health indicators such as availability, latency, error rates, capacity, incident trends, deployment quality, and other reliability metrics. • Create and maintain technical documentation, operational support guides, escalation paths, production readiness checklists, and disaster recovery procedures. • Understand and comply with all company privacy, security, accessibility, change management, and technology standards.

🎯 Requirements

• Bachelor’s degree in Computer Science, Computer Information Systems, Software Engineering, Information Technology, or a related discipline is preferred; equivalent experience or training may be considered. • 3–5+ years of experience in site reliability engineering, DevOps, cloud operations, production support, systems engineering, software engineering, or a related technology operations role. • Experience supporting high-availability web, mobile, commerce, API, integration, or cloud-hosted application environments. • Hands-on experience with monitoring, logging, alerting, incident management, root-cause analysis, CI/CD pipelines, Git-based workflows, and release support. • Experience with cloud platforms, containers, infrastructure automation, scripting, APIs, microservices, content management systems, or restaurant/retail technology environments preferred.

🏖️ Benefits

• Medical, Rx, Dental and Vision Benefits on Day 1 • Life Insurance and Disability Coverage • Paid Vacation/Employee Assistance Program • Business Resource Groups • Tuition Reimbursement • Professional Development • Onboarding, training, and development to help you thrive • Recognition programs and employee events that bring us together • 401k Plan with Company Matching Contributions at 90 days • Employee Stock Purchase Program • 35% Discount on Cracker Barrel Food and Retail items • Exclusive Biscuit Perks like discounts on home, travel, cell phones, and more!

Apply Now

Similar Jobs

🔥 17 hours ago

Fiserv

10,000+ employees

💸 Finance

💳 Fintech

🏦 Banking

Senior DevOps Engineer at Fiserv designing secure, scalable solutions within Microsoft Azure. Collaborating on cloud initiatives and driving operational excellence through Infrastructure as Code and CI/CD automation.

Ansible

Azure

Cloud

Terraform

🔥 20 hours ago

BetterHelp

1 - 10

🏥 Healthcare

⚕️ Healthcare Insurance

🧘 Wellness

Senior DevOps Engineer at BetterHelp responsible for designing and maintaining reliable systems and processes. Collaborating with software engineers and managing production systems.

AWS

Docker

Kubernetes

Linux

Terraform

🕒 Yesterday

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Site Reliability Engineer responsible for automation solutions for global cloud environments. Working with infrastructure teams to ensure operational efficiency and reliability for millions of managed devices.

Ansible

Cloud

Distributed Systems

Linux

RSpec

Ruby

🕒 Yesterday

NBA

11 - 50

🏠 Real Estate

🤝 B2B

Gen AI Security & DevSecOps Engineer responsible for the NBA's software delivery security and AI adoption. Hands-on role securing CI/CD pipelines and cloud environments.

Cloud

Kubernetes

SDLC

🕒 Yesterday

Symbotic

501 - 1000

🏭 Manufacturing

🤖 Artificial Intelligence

🔧 Hardware

Senior Reliability Engineer at Symbotic overseeing RCA investigations for production incidents. Leading problem analysis and improvement initiatives in complex production environments with AI-powered robotics.

Distributed Systems