Senior Reliability Operations Engineer

Job not on LinkedIn

🕒 June 24

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Serve Robotics

Serve Robotics

51 - 200 employees

Founded 2017

📦 Logistics

🍽️ Food & Beverage

🚗 Transport

💰 $30M Venture Round on 2023-08

Logistics • Food & Beverage • Transport

Serve Robotics is an innovative company focused on revolutionizing the delivery industry with its autonomous delivery robots. The company aims to make delivery services more affordable, sustainable, and convenient by using self-driving robots instead of traditional two-ton vehicles for small deliveries like burritos. Through a commercial deal with Uber, Serve Robotics plans to deploy up to 2,000 robots, marking a significant advancement in the autonomous delivery sector.

📋 Description

• Serve as the primary incident lead during your region’s daytime hours, coordinating technical investigations, centralizing communication, and engaging the appropriate engineering and SRE teams when escalation is required. • Respond to escalations from Tier 1 support, using runbooks, metrics, logs, and system diagnostics to investigate and remediate issues or determine when escalation to Tier 3 is necessary. • Develop and update runbooks, workflows, and operational documentation to ensure consistent and reliable responses to recurring issues, collaborating with product teams to expand coverage over time. • Write, maintain, and enhance automation scripts and tools that streamline common remediation steps, improve response times, and reduce manual operational overhead. • Use metrics, logs, and tracing tools (Grafana/Prometheus, GCP Monitoring, OpenTelemetry) to proactively identify problems, validate system behavior, and support continuous improvement of detection mechanisms. • Act as the central point of communication during active incidents, ensuring timely updates and clear routing to the correct product engineering and SRE stakeholders. • Collaborate with reliability and product teams to share insights, recommend improvements, and help refine processes that enhance the stability and operability of our systems. • Participate in a shared weekend on-call rotation to help maintain operational coverage for production systems, responding to incidents and escalations as needed and coordinating with engineering teams when issues arise. • Help establish operational best practices, refine workflows, and prepare the foundation for a broader reliability operations function.

🎯 Requirements

• Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent practical experience. • 5+ years of professional experience in Reliability Operations, Site Reliability Engineering, DevOps, IT Operations, or a related technical support function. • Demonstrated experience owning or participating in Tier 2 or Tier 3 technical investigations, including triage, log analysis, and structured escalation. • Experience supporting distributed systems, cloud-hosted services, or production operational environments. • Hands-on experience participating in incident response processes. • Strong proficiency with Linux, including navigating systems, reviewing logs, and performing diagnostics. • Experience writing, executing, and maintaining runbooks, automations, and operational workflows. • Ability to interpret metrics, logs, and traces using tools such as Grafana/Prometheus, Google Cloud Monitoring, and OpenTelemetry. • Familiarity with modern cloud environments, preferably Google Cloud Platform (GCP), including basic debugging, permissions, and service-level triage. • Ability to investigate and remediate issues following documented procedures, escalating effectively when needed. • Understanding of CI/CD pipelines, deployed application behavior, and operational dependencies across microservices. • Proficiency with Jira or similar platforms for ticketing and structured incident tracking. • Exceptional communication skills, especially during high-pressure incidents where clear, concise updates are critical. • Calm and methodical approach to troubleshooting, prioritization, and decision-making. • Strong collaboration skills when coordinating with product engineering, SRE, and global support teams. • High level of ownership, reliability, and accountability when handling operational responsibilities and incident leadership.

Apply Now

Similar Jobs

🕒 June 8

Remote

501 - 1000

💼 Consulting

📦 Logistics

🏥 Healthcare

Benefits Operations Specialist contributing to global employment compliance at Remote. Supporting employees and clients with HR operations in the APAC region.

🕒 May 17

Zurich Insurance

10,000+ employees

💼 Consulting

🏥 Healthcare

⚖️ Legal

Operations Implementation Analyst focusing on vendor management and operational process improvements for Zurich in Kuala Lumpur. Responsible for data analysis and managing vendor performance.

🕒 March 30

Xsolla

201 - 500

💼 Consulting

📣 Marketing

📦 Logistics

Technical Service Operations Lead coordinating incident response in a dynamic global commerce company. Ensuring platform reliability and uptime for commerce solutions in the video game industry.