Lead Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Ford Motor Company

Ford Motor Company

10,000+ employees

Founded 1903

📦 Logistics

💼 Consulting

📣 Marketing

💰 Post-IPO Debt on 2023-08

Logistics • Consulting • Marketing

Ford Motor Company is a globally renowned automotive company based in the United States, established by Henry Ford. The company is committed to building a better world where every individual has the freedom to move and follow their dreams. Ford is dedicated to innovation, with a focus on services, experiences, and software alongside its traditional vehicle manufacturing. The company is actively involved in sustainability initiatives and aims to meet ambitious environmental targets. Ford values service, community impact, and strives to combine business success with social and environmental responsibility. With a rich history of over 121 years, Ford continues to adapt and lead in the evolving automotive landscape.

📋 Description

• Participate in a 24/7 on-call rotation, responding rapidly to critical incidents and ensuring high availability for the NA eCommerce platform • Diagnose, troubleshoot, and resolve complex production issues as part of incident triage teams, reducing Mean Time to Recovery (MTTR) • Execute and improve operational runbooks and Standard Operating Procedures (SOPs) • Lead and participate in blameless post-mortems and Root Cause Analysis (RCA) sessions • Partner with development and platform teams to architect long-term reliability solutions • Define and track Service Level Indicators (SLIs) and Objectives (SLOs) • Collaborate with Product Owners to establish service levels and manage Error Budgets • Analyze service-health and SLO impacts during monthly release reviews • Use and optimize Dynatrace, GCP Logging, and other observability tools to monitor system health and identify anomalies • Identify observability blind spots and implement solutions for comprehensive system visibility • Manage metrics, dashboards, and alert definitions using Terraform to provision monitoring infrastructure • Design notification strategies and thresholds for KPI/SLO violations • Develop automation scripts, tools, and workflows to reduce toil • Design and implement self-healing mechanisms for common system failures • Implement and manage AI-driven observability solutions for proactive monitoring and predictive maintenance • Coordinate with platform and engineering teams to resolve production bottlenecks and improve processes • Deliver data-driven reports on system health, incident trends, and SRE initiatives to leadership and stakeholders

🎯 Requirements

• Bachelor’s degree in computer science or equivalent • Minimum 3+ years of professional experience in Site Reliability Engineering or DevOps • Deep hands-on experience with Google Cloud Platform, specifically Cloud Run, GKE, and OpenShift • Advanced proficiency in Terraform, including writing reusable modules, managing state, and automating infrastructure provisioning • Experience with comprehensive observability using metrics, events, logs, and traces • Hands-on experience with Dynatrace or similar APM tools such as Datadog or New Relic, including distributed tracing, synthetic monitoring, and code-level profiling • Proficiency in at least one high-level programming language: Java, Node.js, Python, or Go • Ability to read application code to assist with instrumentation and debug complex production issues • Proven experience managing high-severity incidents and the incident lifecycle from triage through mitigation, resolution, and blameless post-mortem/RCA • Availability to participate in a 24/7 on-call rotation

Apply Now

Similar Jobs

🔥 20 hours ago

Weekday

501 - 1000

👗 Fashion

🛒 Retail

🛍️ eCommerce

DevOps Engineer building scalable cloud infrastructure and CI/CD systems for a client. Managing containers, automation, monitoring, and reliability using AWS, Docker, and Kubernetes.

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Grafana

Jenkins

Kafka

Kubernetes

Linux

Microservices

Prometheus

Python

Terraform

Go

🔥 21 hours ago

Weekday (YC W21)

11 - 50

💼 Consulting

👥 HR Tech

☁️ SaaS

DevOps Engineer building cloud infrastructure, CI/CD pipelines, and containerized systems for a Weekday client. Automating provisioning, monitoring reliability, and deployment workflows using AWS, Docker, Kubernetes, and IaC tools.

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Grafana

Jenkins

Kafka

Kubernetes

Linux

Microservices

Prometheus

Python

Terraform

Go

🕒 Yesterday

Empower

10,000+ employees

💸 Finance

💳 Fintech

👥 B2C

Senior DevOps Engineer automating AWS infrastructure and CI/CD for Empower’s financial SaaS products. Improving reliability, security, and developer productivity through tooling, cloud services, and automation.

Angular

Ansible

Apache

AWS

Azure

Chef

Cloud

DNS

Docker

DynamoDB

EC2

Google Cloud Platform

Java

JavaScript

Jenkins

jQuery

Kubernetes

Linux

Maven

MySQL

NGINX

Node.js

Python

React

Ruby

Splunk

Terraform

Go

🕒 Yesterday

Signalmash

51 - 200

💼 Consulting

📦 Logistics

🏥 Healthcare

DevOps Engineer owning Kubernetes, CI/CD, PostgreSQL, observability, and security for Signalmash’s cloud communications platform. Improving reliability, deployment speed, recovery, and infrastructure costs from India.

AWS

Azure

Cloud

Docker

Flux

Google Cloud Platform

Grafana

JavaScript

Kubernetes

Linux

Node.js

Postgres

Prometheus

Python

Shell Scripting

🕒 3 days ago

Neo4j

501 - 1000

☁️ SaaS

🤖 Artificial Intelligence

🏢 Enterprise

Cloud Operations Engineer managing and troubleshooting customer Neo4j database infrastructure. Supporting deployments, monitoring, upgrades, and incidents across AWS, Azure, Google Cloud, virtual, and bare-metal environments.

Ansible

AWS

Azure

Cloud

Google Cloud Platform

Kubernetes

Linux

Neo4j

Prometheus

Terraform