Search Remote Jobs

Senior DevOps Engineer, AI Platform

🕒 September 2

🇨🇦 Canada – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Newfold Digital

Newfold Digital

1001 - 5000 employees

Founded 2021

💼 Consulting

📣 Marketing

📦 Logistics

💰 Venture Round on 2021-01

Consulting • Marketing • Logistics

Newfold Digital is a leading web presence solutions provider serving millions of small-to-medium businesses globally. Through its portfolio of brands, including Bluehost, CrazyDomains, HostGator, Network Solutions, Register. com, Web. com, and many others, Newfold Digital helps customers of all sizes build a digital presence that delivers results and adds value to businesses. With extensive product offerings such as domains, website builders, hosting, security, online marketing, professional website design, and SEO services, along with personalized support, Newfold Digital collaborates with its customers to meet their online presence needs.

📋 Description

• Translate application and platform technical designs into production-ready cloud infrastructure with minimal supervision • Design, provision, operate, and troubleshoot Kubernetes environments, primarily Azure Kubernetes Service and Oracle Kubernetes Engine • Support AI workloads including LiteLLM-based gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines • Design and manage ingress and egress networking, load balancers, DNS, TLS, private connectivity, routing, NAT, firewalls, network policies, and service-to-service communication • Build and operate infrastructure for web applications and backend services, including APIs, databases, caches, queues, scheduled jobs, and event-driven workloads • Build and maintain CI/CD pipelines using Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries • Automate infrastructure provisioning and configuration using Terraform, Helm, Kubernetes manifests, Python, Bash, and related tooling • Implement end-to-end observability using metrics, logs, distributed tracing, dashboards, alerts, health checks, and SLOs • Own production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and infrastructure cost optimization • Create reusable infrastructure patterns that allow engineering teams to launch new services quickly and consistently • Determine required cloud resources, Kubernetes configuration, namespaces, scaling model, and supporting services for new workloads • Configure connectivity, identities, and secrets; provision and operate PostgreSQL, Redis, RabbitMQ, storage, and shared services • Build Jenkins and Bitbucket CI/CD flows for build, test, container publishing, deployment, validation, and rollback • Define observability, health checks, dashboards, alerts, capacity monitoring, and operational runbooks before production launch • Own infrastructure delivery through UAT and production, partnering with architects and engineers on design tradeoffs

🎯 Requirements

• 7 or more years of experience in DevOps, SRE, Platform Engineering, Cloud Infrastructure, or a related role • Strong hands-on experience operating production Kubernetes environments and deep knowledge of networking, scheduling, storage, autoscaling, security, and troubleshooting • Strong Microsoft Azure experience, including AKS, networking, identity, storage, and monitoring • OCI experience is preferred, or demonstrated ability to work across cloud providers • Strong cloud networking knowledge across virtual networks, subnets, routing, NAT, load balancers, private networking, DNS, TLS, firewalls, ingress, and egress • Strong experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code • Proven experience supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures • Hands-on experience with databases, caching, and messaging systems such as PostgreSQL, Redis, RabbitMQ, or equivalent technologies • Experience implementing production observability using OpenTelemetry, Grafana, Prometheus, Sentry, cloud monitoring, or similar tools • Strong Linux, systems, and production troubleshooting skills • Working knowledge of Python, especially backend services built with frameworks such as FastAPI • Familiarity with at least one additional language such as C#, Java, Go, JavaScript, or TypeScript • Understanding of HTTP, HTTPS, DNS, TCP/IP, proxies, authentication, APIs, connection pooling, caching, concurrency, queues, retries, dead letter queues, and asynchronous processing • Ability to read application logs and stack traces and diagnose latency, memory, CPU, connection, and dependency issues

Apply Now

Similar Jobs

🕒 September 1

CruxOCM

11 - 50

⚡ Energy

🏢 Enterprise

☁️ SaaS

Forward Deployed Engineer deploying CruxOCM’s heavy-industry automation software in complex operational technology environments. Integrating customer systems, commissioning solutions, and translating field learnings into product improvements.

🇨🇦 Canada – Remote

💵 $128k - $167k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 31

ComPsych

1001 - 5000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior DevSecOps Engineer modernizing cloud infrastructure, CI/CD, and security for ComPsych’s workplace mental health platform. Guiding AI-built automation across application teams.

🇨🇦 Canada – Remote

💵 $150k - $200k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 29

AbacusNext

201 - 500

⚖️ Legal

💼 Consulting

☁️ SaaS

Senior DevOps Engineer building secure, scalable Azure platforms for CARET’s legal and accounting practice-management software. Leading infrastructure, Kubernetes, CI/CD, security, observability, and reliability initiatives.

🇨🇦 Canada – Remote

💵 $140k - $170k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 28

Penn Interactive

201 - 500

🎲 Gambling

🎮 Gaming

🛍️ eCommerce

Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.

🇨🇦 Canada – Remote

💵 $145k - $193k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 25

JFrog

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior DevOps Engineer helping JFrog customers build CI/CD platforms using JFrog’s liquid software tools. Designing cloud-native pipelines and guiding customers, communities, and internal teams.

🇨🇦 Canada – Remote

💵 $130k - $150k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)