Engineering Manager – Site Reliability

🔥 0 minutes ago

🇨🇦 Canada – Remote

💵 $183k - $193k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Instacart

Instacart

1001 - 5000 employees

Founded 2012

🍽️ Food & Beverage

📦 Logistics

🛍️ eCommerce

💰 $232M Venture Round on 2021-11

Food & Beverage • Logistics • eCommerce

Instacart is a company that offers a flexible approach to work while transforming the grocery industry. It provides an essential service by delivering groceries and household goods to customers' doors in as little as 30 minutes. Instacart offers safe and flexible earning opportunities to personal shoppers and tackles challenges such as rerouting deliveries during snowstorms and connecting customers with coupons and deals. It aims to be the operating system for the grocery industry, thus helping customers save time for other activities. Instacart emphasizes diversity, equity, and belonging in its work culture.

📋 Description

• Lead, mentor, and develop a team of Site Reliability Engineers • Set technical direction and priorities for reliability, scalability, availability, performance, and operational readiness • Partner with engineering, product, security, infrastructure, and other cross-functional teams • Define reliability standards, influence system design, and deliver initiatives improving customer and developer experience • Drive incident management, incident response, post-incident learning, service-level objectives, capacity planning, observability, and continuous risk reduction • Promote automation and self-service tooling to reduce operational toil and improve deployment confidence • Balance near-term operational needs with long-term investments and make tradeoffs in a high-growth environment • Communicate reliability risks, project status, tradeoffs, and decisions to technical and non-technical stakeholders • Make decisions with incomplete information, respond during incidents, and help teams learn from failure

🎯 Requirements

• Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience • Seven or more years of experience in software engineering, infrastructure engineering, Site Reliability Engineering, or a related field • Two or more years of experience managing, mentoring, or leading engineering teams • Professional experience with cloud infrastructure, distributed systems, networking, containers, orchestration platforms, or related production technologies • Experience leading or participating in production incident response, post-incident reviews, reliability improvement initiatives, and operational readiness practices • Experience communicating technical risks, priorities, and tradeoffs to engineering leaders and cross-functional stakeholders • Preferred: experience leading Site Reliability Engineering, platform engineering, infrastructure engineering, or developer productivity teams • Preferred: experience operating highly available services at significant scale and improving service-level objectives, observability, capacity, or disaster recovery capabilities • Preferred: experience with infrastructure as code, continuous delivery, monitoring, logging, tracing, and automated remediation • Preferred: experience building or evolving reliability programs across multiple engineering teams • Preferred: demonstrated ability to create alignment across teams, navigate ambiguity, and turn complex technical challenges into clear, achievable plans • Preferred: leadership approach grounded in empathy, transparency, direct communication, collaboration, and inclusive team development

🏖️ Benefits

• Flexibility to work from home, an office, or a coffee shop • Regular in-person events • New hire equity grant • Annual refresh grants • Competitive compensation and benefits

Apply Now

Similar Jobs

🕒 Yesterday

Rentsync

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental-property software products. Leading incident response, observability, automation, and infrastructure hardening.

🇨🇦 Canada – Remote

💵 $80k - $110k / year

💰 Private equity on 2025-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 Yesterday

commonsku

51 - 200

☁️ SaaS

🤝 B2B

DevOps Engineer II managing AWS infrastructure, IaC, CI/CD, and observability. Supporting commonsku’s SaaS platform for promotional-products distributors.

🕒 5 days ago

Clario

5001 - 10000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior DevOps Engineer advancing AWS cloud infrastructure, CI/CD, and DevSecOps for Clario's clinical research technology. Driving secure, scalable software delivery across global engineering teams.

🕒 5 days ago

Practice EHR

51 - 200

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Intermediate DevOps Engineer building reliable AWS infrastructure and CI/CD automation. Supporting Practice Better’s EHR and practice management platform for health and wellness practitioners.

🕒 September 16

NBCUniversal

10,000+ employees

📱 Media

Principal DevOps Engineer architecting NBCUniversal’s Kubernetes-native platform for broadcast production. Leading Go, AWS, Crossplane, GitOps, networking, security, and observability engineering.