Site Reliability Engineer – Core Streaming

🕒 August 18

🇨🇦 Canada – Remote

💵 $135k - $185k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Yelp

Yelp

1001 - 5000 employees

Founded 2004

🍽️ Food & Beverage

🏨 Hospitality

📣 Marketing

Food & Beverage • Hospitality • Marketing

Yelp is a platform that connects people with great local businesses, allowing users to write and share reviews of a variety of services including restaurants, boutiques, salons, dentists, mechanics, plumbers, and more. With approximately 155 million cumulative reviews contributed by users, Yelp enhances the consumer experience by bringing traditional word-of-mouth recommendations online.

📋 Description

• Own the reliability, scalability, and operational health of Kafka clusters across multi-cloud and hybrid environments • Build and maintain automation for cluster operations, upgrades, capacity scaling, and incident recovery • Partner with engineering teams to enable new streaming use cases, advise on best practices, and ensure data pipeline reliability • Troubleshoot complex issues affecting data flow, performance, or stability • Lead root cause analyses • Execute Kafka version upgrades and platform migrations with minimal disruption to critical services • Participate in on-call rotations using a geographically distributed follow-the-sun model • Drive automation and self-service for deploying, upgrading, and scaling streaming infrastructure

🎯 Requirements

• Solid SRE or infrastructure engineering foundation • Experience with infrastructure-as-code, especially Terraform • Experience with configuration management tools such as Puppet, Ansible, or equivalent • Experience with cloud platforms; AWS preferred • Linux operations experience • Production-level experience with Kafka or similar technologies at scale • Experience with cluster upgrades, migrations, and capacity planning • Programming proficiency in Python, Java, or similar • Strong debugging and systems-thinking skills across distributed systems • Experience with Apache Flink or other stream processing frameworks (nice to have) • Familiarity with Kafka Client APIs, including Producer, Consumer, and Streams (nice to have) • Experience building internal self-service tooling or developer platforms (nice to have) • Experience with incident response and management (nice to have)

🏖️ Benefits

• Fully remote work across Canada • Follow-the-sun on-call model; no one needs to be on-call 24 hours a day • Support from managers, mentors, and teams • Five star benefits (linked in the posting) • Reasonable accommodations for individuals with disabilities in the job application process

Apply Now

Similar Jobs

🕒 August 17

CloudFactory

1001 - 5000

💼 Consulting

📦 Logistics

📣 Marketing

Senior SRE building scalable infrastructure, CI/CD pipelines, and observability systems at CloudFactory. Improving reliability for production environments supporting AI data operations.

🕒 August 14

WinAir

51 - 200

📦 Logistics

💼 Consulting

🏭 Manufacturing

DevOps Specialist automating CI/CD and infrastructure for WinAir’s aviation maintenance software. Improving Jenkins, Ansible, Linux environments, deployments, and monitoring across development and production systems.

🇨🇦 Canada – Remote

💵 $54k - $76k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 14

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

Senior SRE maintaining Kubernetes-based UI and AI service reliability for an enterprise cloud software company. Managing incidents, observability, deployments, and runtime troubleshooting in production.

🇨🇦 Canada – Remote

💰 Private Equity Round on 2020-12

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 13

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

Senior SRE supporting Kubernetes production reliability for an enterprise cloud software company. Troubleshooting distributed services, observability, incidents, CI/CD, Node.js, and JVM/Java runtimes.

🇨🇦 Canada – Remote

💰 Private Equity Round on 2020-12

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 12

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

Site Reliability Engineer operating Kubernetes UI services for a leading enterprise cloud software company. Monitoring reliability, troubleshooting incidents, and supporting production operations.

🇨🇦 Canada – Remote

💰 Private Equity Round on 2020-12

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)