DevOps Engineer, ML Infrastructure

Job not on LinkedIn

🔥 21 minutes ago

🇨🇦 Canada – Remote

💵 $72.1k - $138.9k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 2%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Capgemini

Capgemini

10,000+ employees

Founded 1967

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

Capgemini is a global leader in partnering with businesses to transform and manage their operations by harnessing the power of technology. With expertise across a wide array of industries such as aerospace, automotive, banking, and healthcare, Capgemini provides a constantly evolving portfolio of services to meet the ever-changing needs of their clients. Their offerings include cloud, cybersecurity, data and artificial intelligence, and enterprise management, among others. Capgemini also emphasizes innovation and sustainability, helping companies achieve digital transformation while promoting environmental and social responsibility. Additionally, Capgemini provides career opportunities across various levels and professions, encouraging innovation and diversity in its workforce.

📋 Description

• Build, operate, and evolve large-scale, business-critical infrastructure platforms • Develop and manage cloud-native infrastructure supporting large-scale ML workloads on Kubernetes • Implement and operate batch scheduling solutions such as Volcano and Kueue to optimize GPU and compute resource utilization • Build and maintain CI/CD pipelines for ML services, infrastructure, and platform components • Manage and optimize Kubernetes environments, including production workloads on AWS EKS and other cloud platforms • Lead and support infrastructure modernization and migration initiatives across cloud and platform ecosystems • Partner with Data Scientists, ML Engineers, and Software Engineers to productionize machine learning solutions • Implement observability, monitoring, and alerting for distributed systems and GPU clusters • Ensure platform reliability, security, scalability, and operational excellence • Troubleshoot complex distributed systems and performance bottlenecks across infrastructure and ML workloads • Improve developer productivity and drive innovation through automation and AI-assisted operations

🎯 Requirements

• Strong experience with Kubernetes and large-scale workload orchestration • Hands-on experience with Kubernetes batch schedulers such as Volcano and Kueue • Solid understanding of distributed systems, containerization technologies, cloud-native architectures, and microservices-based platforms • Experience managing workloads on Kubernetes platforms such as AWS EKS • Proven track record delivering and supporting infrastructure migration projects • Strong programming skills in Python or Golang • Experience with Infrastructure-as-Code and deployment tools including Terraform and Helm • Experience designing and maintaining CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, or Azure DevOps • Strong Linux systems administration and troubleshooting skills • Experience building observability solutions using Prometheus and Grafana • Understanding of ML lifecycle management, model deployment, and production operations

🏖️ Benefits

• Paid time off: Vacation (12–25 days depending on grade), company-paid holidays, personal days, and sick leave • Medical, dental, and vision coverage, or provincial healthcare coordination in Canada • Retirement savings plans, including RRSP in Canada • Life and disability insurance • Employee assistance programs • Other benefits provided by local policy and eligibility • Potential eligibility for variable incentives, bonuses, or commissions

Apply Now

Similar Jobs

🔥 23 hours ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud-native tooling and software.

🇨🇦 Canada – Remote

💵 $129.4k - $232.8k / year

💰 Post-IPO Equity on 2001-07

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 4 days ago

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior SRE designing scalable, reliable AWS and Linux infrastructure for Thumbtack’s home-services marketplace. Building platform services, improving availability, and supporting engineering teams.

🇨🇦 Canada – Remote

💵 $180.2k - $233.2k / year

💰 $75M Debt Financing - Thumbtack on 2024-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 4 days ago

Kinaxis

1001 - 5000

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

Cloud Engineer supporting Kinaxis’s AI-powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.

🇨🇦 Canada – Remote

💰 $33M Venture Round on 2000-05

⏰ Full Time

🟢 Junior

🟡 Mid-level

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 5 days ago

Instacart

1001 - 5000

🍽️ Food & Beverage

📦 Logistics

🛍️ eCommerce

Engineering Manager leading Site Reliability Engineers for Instacart’s grocery technology marketplace. Improving platform reliability, scalability, incident response, and operational excellence across engineering teams.

🇨🇦 Canada – Remote

💵 $183k - $193k / year

💰 $232M Venture Round on 2021-11

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 6 days ago

Rentsync

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental-property software products. Leading incident response, observability, automation, and infrastructure hardening.

🇨🇦 Canada – Remote

💵 $80k - $110k / year

💰 Private equity on 2025-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)