Site Reliability Engineer

🕒 il y a 1 mois

🇨🇦 Canada – Télétravail

💵 $125 000 - $250 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

👻 Score fantôme 6%

infoinfo

🗣️🇺🇸🇬🇧 Anglais requis

Firewalls

Kubernetes

Linux

Switching

Postuler Maintenant
Trouver des Emplois à Distance Similaires

📊 Vérifiez votre score de CV pour ce poste

Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

Logo of Boson

Boson

51 - 200 employés

🤖 Intelligence artificielle

🤝 B2B

Artificial Intelligence • B2B

Nous croyons que l'utilisation de l'intelligence artificielle (IA) est nécessaire pour nous conduire à un nouveau niveau de civilisation. Nous nous concentrons sur le développement de produits logiciels de pointe qui peuvent nous propulser vers l'avant.

Description

• Design, operate, and improve reliable infrastructure for AI training and inference workloads • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements • Improve provisioning, configuration management, testing, and deployment automation • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards

🎯 Exigences

• 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role • Strong hands-on expertise in at least one of the following: • Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand • Cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms • Distributed storage, particularly Ceph • GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting • AI training or model-serving infrastructure • Experience operating production systems with a focus on availability, performance, security, and automation • Strong Linux administration and scripting skills • A systematic approach to troubleshooting across multiple layers of a complex system • Clear written and verbal communication skills, including the ability to work effectively with a distributed team

Postuler Maintenant

Emplois Similaires

🕒 il y a 1 mois

SecurityScorecard

501 - 1000

💼 Conseil

🏥 Santé

🛡️ Assurance

Senior Site Reliability Engineer optimizing Kubernetes infrastructure and AI tooling at SecurityScorecard. Driving best practices in automation, observability, and resilience through team collaboration.

🇨🇦 Canada – Télétravail

💵 $197 500 - $225 000 / an

💰 €180 000 000 Series E en 2021-03

⏰ Temps Plein

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 mois

Jonas Software

1001 - 5000

🏗️ Construction

🏥 Santé

🏭 Fabrication

AI-First DevOps Engineer leading AWS infrastructure deployment automation for Computrition. Driving cloud practices and improving DevOps workflows with AI adoption in engineering delivery.

🇨🇦 Canada – Télétravail

💵 $155 000 - $165 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 mois

Carbon60

51 - 200

💼 Conseil

🏥 Santé

📦 Logistique

Managed Services Reliability Engineer supporting Canadian customers’ AWS cloud infrastructure at OpsGuru. Leading incident response, troubleshooting, security, backup, and reliability operations.

🇨🇦 Canada – Télétravail

💵 $140 000 / an

💰 Private Equity Round en 2019-01

⏰ Temps Plein

🟠 Senior

🔴 Expert

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 mois

Smile Digital Health

201 - 500

💼 Conseil

📦 Logistique

📣 Marketing

Site Reliability Engineer responsible for performance and reliability of cloud services at Smile Digital Health. Collaborating with teams to develop and improve performance testing frameworks and systems.

🇨🇦 Canada – Télétravail

💵 $110 000 - $125 000 / an

💰 €30 000 000 Series B en 2023-01

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 mois

Ping Identity

1001 - 5000

💼 Conseil

🏥 Santé

📦 Logistique

Site Reliability Engineer managing AWS accounts and cloud infrastructure deployment. Collaborating with teams to ensure security and efficiency of cloud operations at Ping Identity.

🇨🇦 Canada – Télétravail

💵 $87 000 - $105 000 / an

💰 €35 000 000 Series F - Ping Identity en 2014-09

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

Ansible

AWS

Chef

Cloud

Linux

Puppet

Python

Ruby

SaltStack

Terraform

Unix

Go