
1001 - 5000 employees
Founded 2004
🍽️ Food & Beverage
🏨 Hospitality
📣 Marketing
Food & Beverage • Hospitality • Marketing
Yelp is a platform that connects people with great local businesses, allowing users to write and share reviews of a variety of services including restaurants, boutiques, salons, dentists, mechanics, plumbers, and more. With approximately 155 million cumulative reviews contributed by users, Yelp enhances the consumer experience by bringing traditional word-of-mouth recommendations online.
🕒 August 18
🇨🇦 Canada – Remote
💵 $135k - $185k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
👻 Ghost score 0%
Improve your chances of getting an interview by checking your resume score before you apply.

1001 - 5000 employees
Founded 2004
🍽️ Food & Beverage
🏨 Hospitality
📣 Marketing
Food & Beverage • Hospitality • Marketing
Yelp is a platform that connects people with great local businesses, allowing users to write and share reviews of a variety of services including restaurants, boutiques, salons, dentists, mechanics, plumbers, and more. With approximately 155 million cumulative reviews contributed by users, Yelp enhances the consumer experience by bringing traditional word-of-mouth recommendations online.
• Own the reliability, scalability, and operational health of Kafka clusters across multi-cloud and hybrid environments • Build and maintain automation for cluster operations, upgrades, capacity scaling, and incident recovery • Partner with engineering teams to enable new streaming use cases, advise on best practices, and ensure data pipeline reliability • Troubleshoot complex issues affecting data flow, performance, or stability • Lead root cause analyses • Execute Kafka version upgrades and platform migrations with minimal disruption to critical services • Participate in on-call rotations using a geographically distributed follow-the-sun model • Drive automation and self-service for deploying, upgrading, and scaling streaming infrastructure
• Solid SRE or infrastructure engineering foundation • Experience with infrastructure-as-code, especially Terraform • Experience with configuration management tools such as Puppet, Ansible, or equivalent • Experience with cloud platforms; AWS preferred • Linux operations experience • Production-level experience with Kafka or similar technologies at scale • Experience with cluster upgrades, migrations, and capacity planning • Programming proficiency in Python, Java, or similar • Strong debugging and systems-thinking skills across distributed systems • Experience with Apache Flink or other stream processing frameworks (nice to have) • Familiarity with Kafka Client APIs, including Producer, Consumer, and Streams (nice to have) • Experience building internal self-service tooling or developer platforms (nice to have) • Experience with incident response and management (nice to have)
• Fully remote work across Canada • Follow-the-sun on-call model; no one needs to be on-call 24 hours a day • Support from managers, mentors, and teams • Five star benefits (linked in the posting) • Reasonable accommodations for individuals with disabilities in the job application process
Apply Now🕒 August 17
Senior SRE building scalable infrastructure, CI/CD pipelines, and observability systems at CloudFactory. Improving reliability for production environments supporting AI data operations.
🕒 August 14
DevOps Specialist automating CI/CD and infrastructure for WinAir’s aviation maintenance software. Improving Jenkins, Ansible, Linux environments, deployments, and monitoring across development and production systems.
🇨🇦 Canada – Remote
💵 $54k - $76k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 August 14
Senior SRE maintaining Kubernetes-based UI and AI service reliability for an enterprise cloud software company. Managing incidents, observability, deployments, and runtime troubleshooting in production.
🇨🇦 Canada – Remote
💰 Private Equity Round on 2020-12
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 August 13
Senior SRE supporting Kubernetes production reliability for an enterprise cloud software company. Troubleshooting distributed services, observability, incidents, CI/CD, Node.js, and JVM/Java runtimes.
🇨🇦 Canada – Remote
💰 Private Equity Round on 2020-12
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 August 12
Site Reliability Engineer operating Kubernetes UI services for a leading enterprise cloud software company. Monitoring reliability, troubleshooting incidents, and supporting production operations.
🇨🇦 Canada – Remote
💰 Private Equity Round on 2020-12
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)