Lead Site Reliability Engineer

🕒 il y a 4 mois

🇨🇦 Canada – Télétravail

💵 $154 000 - $200 000 / an

⏰ Temps Plein

🟠 Senior

⛑ Ingénieur DevOps & SRE

👻 Score fantôme 27%

infoinfo

🗣️🇺🇸🇬🇧 Anglais requis

Postuler Maintenant
Trouver des Emplois à Distance Similaires

📊 Vérifiez votre score de CV pour ce poste

Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

Logo of Movable Ink

Movable Ink

501 - 1000 employés

📣 Marketing

✈️ Tourisme

🏨 Hôtellerie

💰 €55 000 000 Series D en 2022-04

Marketing • Travel • Hospitality

Movable Ink est un leader en marketing personnalisé, utilisant l'IA et l'automatisation pour offrir un contenu sur mesure aux clients à travers divers points de contact, tels que les e-mails et les mobiles. Leurs solutions, comprenant le Movable Ink Studio et Da Vinci, renforcent les secteurs de la vente au détail, des services financiers, des organisations médiatiques et des voyages & hôtellerie en transformant les données en expériences client uniques en temps réel. L'entreprise collabore avec un vaste réseau de partenaires stratégiques, améliorant les technologies marketing existantes pour atteindre des résultats remarquables en matière de revenus et d'engagement client. L'expertise de Movable Ink est validée par des études, telles que l'étude Total Economic Impact™ de Forrester Consulting, qui met en évidence des avantages financiers significatifs pour les utilisateurs. Grâce à l'utilisation innovante de l'IA, Movable Ink stimule la personnalisation au prochain niveau et la performance marketing.

Description

• Define and drive the automation strategy for infrastructure tooling, establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents • Own the design, reliability and evolution of core platform applications, mentoring team members on best practices and ensuring systems meet long-term business objectives • Architect and lead the logging platform strategy, driving its design and balancing availability, retention and cost optimization • Establish capacity planning and performance management frameworks, proactively identifying scaling opportunities and guiding teams through complex troubleshooting scenarios • Lead cross-functional reliability initiatives with SRE and service engineering teams, influencing architectural decisions and championing practices that ensure resilient service delivery • Demonstrate a high level of autonomy in anticipating, identifying, and addressing systemic weaknesses and opportunities for platform improvement without direct supervision.

🎯 Exigences

• Proven track record in Site Reliability or Software Engineering, designing, building, and owning scalable, resilient services with a focus on long-term reliability strategy • Deep expertise in architecting and operating complex distributed systems such as Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB/Cassandra, with the ability to guide teams through distributed system challenges • Designing and owning automation strategies to manage services at scale, with expertise in establishing performance analysis frameworks and mentoring others on diagnostics and resolution • Deep, hands-on experience (6+ years) in Site Reliability or Software Engineering, specifically leading and shaping multi-cloud architecture and strategy (AWS and GCP). • Experience architecting and leading large-scale observability platforms, including defining observability standards and SLO frameworks. We use Prometheus and Thanos with Grafana Alloy, Loki and Tempo • Experience leading on-call excellence, including driving improvements to monitoring and alerting strategies, automating runbooks and mentoring team members on incident response best practices. Every member of the SRE team does a week long on-call rotation • Expert-level proficiency with infrastructure as code, including defining IaC standards and patterns across teams. We use Terraform and Chef • Advanced Kubernetes expertise, including cluster architecture design, multi-tenancy strategies, and guiding teams on container orchestration best practices. We use EKS and GKE • Proficiency in multiple programming languages with the ability to design and review code that meets reliability standards. We use NodeJS, Golang, Ruby, Python and shell scripting • Advanced Linux systems expertise, with the ability to diagnose complex system-level issues and mentor others on performance tuning and troubleshooting.

🏖️ Avantages

• full range of medical • financial • other benefits

Postuler Maintenant

Emplois Similaires

🕒 il y a 4 mois

Tecsys Inc.

501 - 1000

🏥 Santé

☁️ SaaS

📦 Logistique

Ingénieur fiabilité des infrastructures pour soutenir les services SaaS critiques. Collaborer, innover et optimiser la fiabilité et la performance des systèmes cloud sur AWS et Kubernetes.

🇨🇦 Canada – Télétravail

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🕒 il y a 4 mois

Sparrow Connected

11 - 50

💼 Conseil

🏥 Santé

📦 Logistique

DevOps Specialist taking over build, release, and environments for Sparrow’s product team. Leading DevOps practices while collaborating with CTO and senior developers in an agile setting.

🇨🇦 Canada – Télétravail

💰 Seed Round en 2021-03

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 4 mois

My Personal Recruiter

11 - 50

👥 RH Tech

🎯 Recrutement

👥 B2C

DevOps Engineer supporting NY operations from Canada for a global software services provider. Focused on developing and deploying services in a collaborative environment with various technical stacks.

🇨🇦 Canada – Télétravail

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 5 mois

Tyk

51 - 200

🔌 API

☁️ SaaS

🏢 Entreprise

Site Reliability Engineer maintaining and improving Tyk's API Management platform with a focus on operational excellence. Working in a fully remote environment supporting a global, distributed team.

🇨🇦 Canada – Télétravail

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 5 mois

Cority

201 - 500

🏥 Santé

📦 Logistique

💼 Conseil

Sr. DevOps Engineer for Cority working on deployment and operation of systems. Collaborating to deliver automated cloud infrastructures and continuous delivery processes in a remote Canada role.

🇨🇦 Canada – Télétravail

⏰ Temps Plein

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis