Site Reliability Engineering Team Lead – Principal SRE

Job not on LinkedIn

🔥 6 minutes ago

🌐 Canada, United States – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cerence Inc.

Cerence Inc.

1001 - 5000 employees

Founded 2019

💼 Consulting

📦 Logistics

🏭 Manufacturing

💰 Grant on 2020-12

Consulting • Logistics • Manufacturing

Cerence Inc. is a global company focused on providing AI-powered solutions, particularly in the automotive industry. They specialize in conversational and generative AI technologies that create intelligent, natural, and personalized interactions between humans and vehicles. With innovations like their proprietary automotive large language models, Cerence enhances user experiences across various forms of transport including cars, two-wheelers, and trucks. The company has over 500 million vehicles shipped with its AI technology, serving more than 80 OEMs and Tier 1 customers worldwide. Cerence is dedicated to continuous advancements in AI, aiming to revolutionize in-car user experiences through fast delivery and seamless integration of their solutions.

📋 Description

• Lead Cerence's Site Reliability Engineering team and own the reliability, availability, and operational health of its cloud-native automotive AI platform • Help select, mentor, and technically develop the team across multiple locations • Set technical direction and priorities and contribute performance and growth input to team managers • Design and maintain a sustainable on-call rotation and monitor page load and team health • Own and drive the reliability roadmap across a 2–3 quarter horizon • Define and govern SLI/SLO/SLA frameworks for customer program availability targets up to 99.95% • Serve as Tier 2 technical escalation point for major incidents in partnership with the Global Operations Center • Champion blameless postmortem culture and ensure actionable outcomes • Lead and improve Production Readiness / NFR reviews with development teams • Contribute to root cause analysis and own systemic improvements • Approve high-risk and out-of-window production changes • Set strategic direction for metrics, dashboards, alerting, escalation, and automation • Drive CI/CD automation pipelines for service deployments, rollbacks, and operational tasks • Partner with DevOps and platform teams to evolve shared infrastructure • Embed reliability into the SDLC through collaboration with development managers and architects • Participate in service reliability consulting and architectural reviews • Communicate reliability posture and risk to technical and non-technical stakeholders

🎯 Requirements

• 8+ years of hands-on experience in site reliability, DevOps, or cloud platform roles, including time leading a team or owning a function • A track record of setting technical direction and holding standards across a team — with or without formal authority • Hands-on experience with container orchestration frameworks (Kubernetes, Docker, Istio) • Experience with public cloud platforms (Azure primarily; AWS and Google Cloud) • Familiarity with observability tooling — metrics pipelines, dashboarding, and alerting (e.g., Zabbix, Prometheus, Grafana) • Experience with CI/CD pipelines and infrastructure-as-code practices (e.g., Terraform, Flux) • Proficiency in at least one scripting or programming language (Python, Go, Shell, etc.) • Strong UNIX/Linux background, including system configuration, performance debugging, and network fundamentals (Layer 4/5, DNS, HTTP/S, TLS) • Excellent written and verbal communication skills in English • Previous site reliability leadership experience • Experience leading distributed or multi-site technical teams • Background in high-availability service design (redundancy, failover, blast radius) • Experience with log aggregation and analytics platforms (Loki, Thanos) • Familiarity with ITSM and project tooling (Jira, Confluence) • Experience in automotive, embedded, or latency-sensitive production environments

🏖️ Benefits

• Annual bonus opportunity • Insurance coverage (medical, dental, vision, life, and disability) • Paid time off • Paid holidays • Company contribution to the RRSP (Registered Retirement Savings Plan) • Equity awards for certain positions and levels • Remote and/or hybrid work available depending on the position

Apply Now

Similar Jobs

🕒 Yesterday

Case IQ

51 - 200

💼 Consulting

🏥 Healthcare

⚖️ Legal

DevOps Engineer optimizing Azure and AWS infrastructure for Case IQ’s governance, risk, and compliance software. Automating deployments, improving reliability, security, observability, and AI-enabled operations.

🕒 Yesterday

GBL

11 - 50

🎮 Gaming

💳 Fintech

DevOps and Golang Developer building Golang microservices and automated infrastructure. Scaling secure systems for a crypto-first international online casino.

🕒 Yesterday

BMO U.S.

5001 - 10000

🛡️ Insurance

💼 Consulting

📦 Logistics

Sr. SRE Engineer ensuring BMO’s Amazon Connect CCaaS platform reliability, security, and availability. Automating operations, observability, incident response, and resilient service delivery.

🕒 5 days ago

PATH

1001 - 5000

DevOps Specialist automating deployments, testing, and operations for Alberta government digital transformation projects. Supporting CI/CD, observability, production operations, and quality delivery.

🕒 6 days ago

Sureify

201 - 500

💳 Fintech

🛡️ Insurance

☁️ SaaS

Senior DevOps Engineer scaling AWS infrastructure for Sureify’s insurance SaaS platform. Automating deployments, reliability, monitoring, and security across the Americas.