Lead Site Reliability Engineer

🕒 June 19

🇨🇦 Canada – Remote

💵 $104k - $143k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 20%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Vista

Vista

5001 - 10000 employees

Founded 1995

📣 Marketing

🛍️ eCommerce

💰 $40M Venture Round on 2003-07

Marketing • eCommerce • Design/UX

Vista is a remote-first company dedicated to empowering small businesses by providing exceptional services in technology, design, marketing, and customer care. With a global workforce of over 6,500 team members across 17 countries, Vista emphasizes collaboration, inclusivity, and sustainability. Their commitment to innovation enables small businesses to thrive, while the company's culture fosters personal and professional growth for its employees.

📋 Description

• Identify recurring technical root causes behind customer-impacting incidents • Prioritise reliability improvements that reduce Mean Time to Detect and Mean Time to Resolve and prevent future incidents • Define engineering interventions including safe deployment defaults, secret and credential rotation, resilience patterns, observability and alerting standards, contract testing, and infrastructure pre-flight validation • Influence owning teams to implement reliability improvements • Provide hands-on support through Merge Requests, pairing, code review, and technical support • Disseminate and evangelise best practices across the organisation • Lead technical discussions in post-incident reviews and operational forums • Identify missing monitoring, untested failure modes, incomplete rollback strategies, and dependency risks • Help grow the Incident Response Team's engineering practice through pairing and internal learning sessions • Partner with Platform Engineering, Deployment Platform, Developer Productivity, and senior engineers across Vista • Review proposals affecting reliability, deployment safety, observability, or operational readiness • Participate in the team's on-call rotation and contribute incident leadership when required

🎯 Requirements

• 5 or more years of hands-on Site Reliability, Platform, or Infrastructure Engineering experience in a large-scale, distributed production environment • Proficiency in at least one programming language, such as Python, Go, TypeScript, or Java • Track record of code shipped to production • Experience driving adoption of a reliability or platform pattern across teams that did not report to you, with measurable outcomes • Strong systems thinking and bias toward simple solutions • Hands-on experience with at least one major cloud platform: AWS, Google Cloud, or Azure • Experience with an observability platform, such as New Relic, Datadog, or Grafana • Experience defining and operating against Service Level Objectives • Experience with continuous integration and deployment pipelines • Experience with infrastructure-as-code, such as AWS CDK or Pulumi • Hands-on exposure to Artificial Intelligence and Large Language Model tooling in an engineering context • Nice to have: mentoring, internal learning programmes, chaos engineering, GameDays, failure-injection programmes, globally distributed engineering organisations, incident-management and service-catalogue tooling, internal developer platforms, paved-road tooling, or golden-path patterns

🏖️ Benefits

• Comprehensive benefits package • Health programs • Wealth programs • Wellness programs • Long-term equity incentives, subject to eligibility • Remote-first work environment

Apply Now

Similar Jobs

🕒 June 6

Capgemini

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Software Change Management Consultant supporting application migration projects using IBM’s DBB/Git/IDD Solutions. Guiding clients through the conversion process and providing migration expertise and training.

Groovy

🕒 April 16

Workiy Inc.

11 - 50

💼 Consulting

📣 Marketing

🛍️ eCommerce

Senior Salesforce DevOps & Release Manager at Workiy overseeing enterprise Salesforce deployments and CI/CD processes across multiple environments. Driving best practices in release management while collaborating with cross-functional teams.

🕒 April 16

Workiy Inc.

11 - 50

💼 Consulting

📣 Marketing

🛍️ eCommerce

Senior Salesforce DevOps Consultant driving DevOps best practices and managing deployment strategies for an IT solutions company. Supporting seamless Salesforce releases across multiple environments and business teams.

🕒 April 15

Movable Ink

501 - 1000

📣 Marketing

✈️ Travel

🏨 Hospitality

Lead Site Reliability Engineer ensuring scalable, resilient services for Movable Ink at high volume content platform. Design and drive automation strategies while mentoring engineering teams.

🇨🇦 Canada – Remote

💵 $154k - $200k / year

💰 $55M Series D on 2022-04

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 April 1

Tecsys Inc.

501 - 1000

🏥 Healthcare

☁️ SaaS

📦 Logistics

IngÊnieur fiabilitÊ des infrastructures pour soutenir les services SaaS critiques. Collaborer, innover et optimiser la fiabilitÊ et la performance des systèmes cloud sur AWS et Kubernetes.

🗣️🇫🇷 French Required