Site Reliability Engineer

Job not on LinkedIn

🕒 July 14

🇨🇦 Canada – Remote

💵 $125k - $250k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 3%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Boson

Boson

51 - 200 employees

🤖 Artificial Intelligence

🤝 B2B

Artificial Intelligence • B2B

Boson is a software development company that believes in the transformative power of Artificial Intelligence (AI) to advance civilization. They focus on creating cutting-edge software products that push technological boundaries and aim to integrate AI seamlessly into everyday life.

📋 Description

• Design, operate, and improve reliable infrastructure for AI training and inference workloads • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements • Improve provisioning, configuration management, testing, and deployment automation • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards

🎯 Requirements

• 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role • Strong hands-on expertise in at least one of the following: • Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand • Cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms • Distributed storage, particularly Ceph • GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting • AI training or model-serving infrastructure • Experience operating production systems with a focus on availability, performance, security, and automation • Strong Linux administration and scripting skills • A systematic approach to troubleshooting across multiple layers of a complex system • Clear written and verbal communication skills, including the ability to work effectively with a distributed team

Apply Now

Similar Jobs

🕒 July 14

SecurityScorecard

501 - 1000

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Senior Site Reliability Engineer optimizing Kubernetes infrastructure and AI tooling at SecurityScorecard. Driving best practices in automation, observability, and resilience through team collaboration.

🇨🇦 Canada – Remote

💵 $197.5k - $225k / year

💰 $180M Series E on 2021-03

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 13

Jonas Software

1001 - 5000

🏗️ Construction

🏥 Healthcare

🏭 Manufacturing

AI-First DevOps Engineer leading AWS infrastructure deployment automation for Computrition. Driving cloud practices and improving DevOps workflows with AI adoption in engineering delivery.

🇨🇦 Canada – Remote

💵 $155k - $165k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 10

Carbon60

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Managed Services Reliability Engineer supporting Canadian customers’ AWS cloud infrastructure at OpsGuru. Leading incident response, troubleshooting, security, backup, and reliability operations.

🇨🇦 Canada – Remote

💵 $140k / year

💰 Private Equity Round on 2019-01

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 10

Smile Digital Health

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Site Reliability Engineer responsible for performance and reliability of cloud services at Smile Digital Health. Collaborating with teams to develop and improve performance testing frameworks and systems.

🇨🇦 Canada – Remote

💵 $110k - $125k / year

💰 $30M Series B on 2023-01

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 8

Ping Identity

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer managing AWS accounts and cloud infrastructure deployment. Collaborating with teams to ensure security and efficiency of cloud operations at Ping Identity.

🇨🇦 Canada – Remote

💵 $87k - $105k / year

💰 $35M Series F - Ping Identity on 2014-09

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)