Site Reliability Engineer, Core Streaming

Job not on LinkedIn

🔥 1 hour ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Yelp

Yelp

1001 - 5000 employees

Founded 2004

🍽️ Food & Beverage

🏨 Hospitality

📣 Marketing

Food & Beverage • Hospitality • Marketing

Yelp is a platform that connects people with great local businesses, allowing users to write and share reviews of a variety of services including restaurants, boutiques, salons, dentists, mechanics, plumbers, and more. With approximately 155 million cumulative reviews contributed by users, Yelp enhances the consumer experience by bringing traditional word-of-mouth recommendations online.

📋 Description

• Own the reliability, scalability, and operational health of Kafka clusters across multi-cloud and hybrid environments • Build and maintain automation for cluster operations, upgrades, capacity scaling, and incident recovery • Partner with engineering teams to enable new streaming use cases, advise on best practices, and ensure data pipeline reliability • Troubleshoot complex issues affecting data flow, performance, or stability, and lead root cause analyses • Execute Kafka version upgrades and platform migrations with minimal disruption to critical services • Participate in on-call rotations using a geographically distributed follow-the-sun model • Drive automation and self-service for deploying, upgrading, and scaling streaming infrastructure

🎯 Requirements

• Solid SRE or infrastructure engineering foundation • Experience with infrastructure-as-code, especially Terraform • Experience with configuration management using Puppet, Ansible, or equivalent • Experience with cloud platforms, AWS preferred • Experience with Linux operations • Production-level experience with Kafka or similar technologies at scale • Experience with cluster upgrades, migrations, and capacity planning • Programming proficiency in Python, Java, or similar • Strong debugging and systems-thinking skills • Ability to trace data-flow issues end-to-end across distributed systems • Experience with Apache Flink or other stream-processing frameworks preferred • Familiarity with Kafka Client APIs, including Producer, Consumer, and Streams preferred • Experience building internal self-service tooling or developer platforms preferred • Experience with incident response and management preferred

🏖️ Benefits

• Fully remote work across Canada • Five star benefits • Follow-the-sun on-call model, so no one needs to be on-call 24 hours a day • Reasonable accommodations for individuals with disabilities • Equal opportunity employment • Opportunities for meaningful ownership and innovation • Support from managers, mentors, and teams

Apply Now

Similar Jobs

🕒 Yesterday

GE Vernova

10,000+ employees

💼 Consulting

📦 Logistics

🏭 Manufacturing

Senior reliability engineer strengthening embedded protection, control, and software products for GE Vernova’s utility-scale grid modernization solutions. Leading testing, failure analysis, and reliability initiatives.

🕒 2 days ago

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

Site Reliability Engineer maintaining Kubernetes UI services for an enterprise cloud software client. Monitoring production health, troubleshooting incidents, and supporting reliable cloud-native operations.

🇨🇦 Canada – Remote

💰 Private Equity Round on 2020-12

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 5 days ago

Mirantis

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Site Reliability Engineer at Mirantis, contributing to cloud-based AI solutions using Kubernetes. Focused on deploying AI infrastructure and ensuring system reliability and performance.

🕒 6 days ago

PerfectServe

201 - 500

🏥 Healthcare

⚕️ Healthcare Insurance

☁️ SaaS

Core role deploying AI products to new customers in healthcare setting. Collaborate with teams to optimize deployments and improve AI offerings.

🇨🇦 Canada – Remote

💵 $165k - $195k / year

💰 Private Equity Round on 2018-05

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 6 days ago

Mirantis

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Site Reliability Engineer deploying AI infrastructure for cloud technologies. Collaborate with international teams on AI-driven automation and high-performance systems using Kubernetes.