Cloud Engineer, Azure Platform Engineer

đŸ”„ 0 minutes ago

🇹🇩 Canada – Remote

đŸ’” $115k - $130k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

☁ Cloud Engineer

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Smile Digital Health

Smile Digital Health

201 - 500 employees

Founded 2016

đŸ’Œ Consulting

📩 Logistics

📣 Marketing

💰 $30M Series B on 2023-01

Consulting ‱ Logistics ‱ Marketing

Smile Digital Health is a company specializing in healthcare data interoperability and health IT solutions. Their Smile Health Data Platform enables efficient data ingestion and transformation into FHIR resources, enhancing clinical decision-making and business intelligence. Focused on compliance and security, Smile Digital Health offers services such as prior authorization automation, standardizing data exchange, and supporting healthcare providers, payers, and health information exchanges with customizable and scalable solutions.

📋 Description

‱ Act as the subject-matter expert (SME) for Kubernetes deployments, troubleshooting, and production issues across all environments. ‱ Design and deploy Azure Kubernetes Service (AKS) clusters with private cluster configurations, managed identities, and RBAC. ‱ Own the health, scaling, and lifecycle management of production AKS clusters, including upgrades, node pool management, autoscaling, and capacity planning. ‱ Configure AKS networking, including Azure CNI, internal load balancers, and ingress controllers (NGINX, Traefik). ‱ Design and maintain integrations between AKS with Azure Container Registry (ACR), Key Vault via CSI driver, Azure Monitor for containers, Azure SQL, Kafka/Event Hubs, Azure Storage, and other client-facing dependencies (DNS resolution, firewall rules, private endpoints, and service connectivity). ‱ Manage containerized application deployments using Docker and Helm; maintain reusable chart and templating standards, namespaces, resource quotas, and Azure Policy for AKS. ‱ Harden AKS environments through policy enforcement, network policies, and image scanning. ‱ Own container and cluster vulnerability management: scanning, triage, prioritization, and remediation coordination with engineering teams. ‱ Contribute to Terraform-based infrastructure as code for provisioning and managing Azure resources. ‱ Support Azure DevOps (or equivalent) CI/CD pipelines, including GitOps workflows (Flux/ArgoCD), to reduce deployment risk and improve release velocity. ‱ Design, implement, and manage observability across the Grafana stack (Prometheus, Loki, Tempo) and Azure-native tooling (Azure Monitor, Log Analytics Workspace, Application Insights) for critical systems; define SLIs/SLOs and tune alerting to reduce noise. ‱ Collaborate with performance engineering and application teams to identify, diagnose, and resolve performance bottlenecks spanning the AKS platform and its dependent services (database, messaging, network) to right-size node pools and workloads based on observed performance and utilization trends. ‱ Contributing to Disaster Recovery and Business Continuity Planning (DR/BCP) procedures for AKS-hosted workloads, including cross-team failover drill participation and RTO/RPO validation. ‱ Provide escalation support for production Kubernetes and infrastructure incidents; participate in on-call rotation, lead root-cause analysis, and drive preventative follow-up actions. ‱ Document runbooks and post-incident reviews, and operational knowledge; maintain a living knowledge base to reduce tribal knowledge.

🎯 Requirements

‱ 5+ years of hands-on Kubernetes experience, including at least 2 years running AKS in production. ‱ Strong knowledge of Kubernetes internals scheduling, networking, storage, RBAC. ‱ Proficiency in Azure CNI networking and AKS private cluster configuration. ‱ Hands-on experience integrating AKS with Azure PaaS services (ACR, AKV, Azure SQL, Kafka and Managed Identities) and troubleshooting network-layer dependencies (DNS, firewall, private endpoints). ‱ Experience with Helm and GitOps workflows (Flux/ArgoCD). ‱ Working knowledge of Terraform for infrastructure as code. ‱ Working knowledge of Azure DevOps or similar CI/CD tooling. ‱ Hands-on experience implementing and operating observability platforms: Grafana stack (Prometheus, Loki, Tempo) and Azure-native monitoring (Azure Monitor, Log Analytics, Application Insights); experience defining SLIs/SLOs for production systems. ‱ Practical experience with container/cluster vulnerability management and remediation workflows. ‱ Scripting skills in Bash, Python. ‱ Certified Kubernetes Administrator (CKA), Or CKAD certification and Microsoft Certified: Azure Administrator (AZ-104) required. ‱ Excellent written and verbal communication skills; able to convey technical issues clearly to both technical and non-technical stakeholders.

đŸ–ïž Benefits

‱ Remote Work Environment ‱ Flexible Time Away From Work Policy including PTO, Personal and Sick Days ‱ Competitive Salary and Health/Medical Benefits ‱ RRSP/TFSA/401K Employee Contribution ‱ Life and Disability ‱ Employee Assistance Program ‱ FHIR Study Program and Skillsoft Learning ‱ Super HAPI Fun Club

Apply Now

Similar Jobs

🕒 Yesterday

DoiT International

201 - 500

☁ SaaS

đŸ’Œ Consulting

Senior Cloud Architect leading the design and implementation of AI and ML solutions at DoiT. Working remotely from Canada and collaborating globally with customers and teams.

🕒 2 days ago

Redpanda Data

51 - 200

🏱 Enterprise

☁ SaaS

đŸ€ B2B

Senior Software Engineer designing and operating cloud services across AWS, GCP, and Azure for Redpanda. Leading technical direction and ensuring reliability and performance in developed systems.

🇹🇩 Canada – Remote

đŸ’” $205k - $240k / year

💰 $100M Series D - Redpanda Data on 2025-04

⏰ Full Time

🟠 Senior

☁ Cloud Engineer

🕒 2 days ago

DoiT International

201 - 500

☁ SaaS

đŸ’Œ Consulting

Senior Cloud Architect in AI-focused role for DoiT, leading design and implementation of Generative AI solutions on AWS. Collaborating with global teams and influencing product strategies.

🕒 6 days ago

CoVet

11 - 50

đŸ’Œ Consulting

📩 Logistics

đŸ„ Healthcare

Full-Stack Developer focused on mobile and cloud solutions for veterinary professionals with a strong emphasis on Flutter. Collaborate across the stack to enhance workflows and improve patient care.

🇹🇩 Canada – Remote

đŸ’” $110k - $130k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

☁ Cloud Engineer

🕒 July 22

Carbon60

51 - 200

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

Senior Cloud Software Architect developing modern cloud-based applications at OpsGuru. Collaborating with clients and teams to deliver technology solutions for digital transformation.

🇹🇩 Canada – Remote

đŸ’” $160k - $180k / year

💰 Private Equity Round on 2019-01

⏰ Full Time

🟠 Senior

☁ Cloud Engineer