
11 - 50 employees
Founded 2024
đď¸ Construction
đ Manufacturing
đŚ Logistics
Construction ⢠Manufacturing ⢠Logistics
Nexxa. ai is an industrial AI company building agentic automation for industrial operations. Its platform learns from engineers and employees to take over complex, repetitive workflows, autonomously setting goals and execution criteria to increase productivity and resilience. Delivered as an integrated solution that works with existing technology stacks, Nexxa. ai positions itself as a SaaS-like enterprise tool that generates immediate ROI for operations teams.
đĽ 0 minutes ago
đ¨đŚ Canada â Remote
â° Full Time
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đť Ghost score 18%
Amazon Redshift
AWS
Azure
BigQuery
Cloud
Distributed Systems
Google Cloud Platform
Grafana
Jenkins
Kubernetes
Prometheus
Python
Terraform
Go
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
Founded 2024
đď¸ Construction
đ Manufacturing
đŚ Logistics
Construction ⢠Manufacturing ⢠Logistics
Nexxa. ai is an industrial AI company building agentic automation for industrial operations. Its platform learns from engineers and employees to take over complex, repetitive workflows, autonomously setting goals and execution criteria to increase productivity and resilience. Delivered as an integrated solution that works with existing technology stacks, Nexxa. ai positions itself as a SaaS-like enterprise tool that generates immediate ROI for operations teams.
⢠Own and evolve Nexxa's core infrastructure, including compute, networking, storage, and deployment systems, end-to-end ⢠Design and operate CI/CD pipelines supporting safe iteration across AI, data, and product engineering teams ⢠Build and maintain infrastructure-as-code for reproducible, auditable cloud and on-prem/edge environments ⢠Architect and manage Kubernetes platforms for training, inference, and application workloads, including GPU scheduling and autoscaling ⢠Support data warehouses, lakehouse architectures, feature stores, embedding indices, retrieval pipelines, and model training, evaluation, and serving infrastructure ⢠Define and drive observability practices covering metrics, logging, tracing, and alerting across distributed systems ⢠Establish reliability practices including SLOs/SLIs, incident response, postmortems, and on-call rotations ⢠Design security and compliance across cloud infrastructure, secrets management, and access control ⢠Make pragmatic tradeoffs among cost, latency, reliability, and developer velocity ⢠Collaborate with engineering leadership on infrastructure roadmap and platform strategy ⢠Mentor engineers on infrastructure best practices and promote operational excellence
⢠6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles ⢠Deep hands-on experience with AWS, GCP, or Azure at production scale ⢠Kubernetes production experience, including GPU workload scheduling ⢠Infrastructure-as-code experience with Terraform, Pulumi, or equivalent ⢠CI/CD systems experience, including GitHub Actions, GitLab CI, CircleCI, Jenkins, or ArgoCD ⢠Experience designing and operating observability stacks such as Prometheus, Grafana, Datadog, or OpenTelemetry ⢠Experience supporting ML/AI infrastructure is a strong plus ⢠Excellent scripting/programming skills in Python, Go, or Bash ⢠Proven ability to independently scope and lead infrastructure projects from design through production rollout ⢠Strong incident management instincts and ability to lead through outages and drive root-cause analysis ⢠Experience with cloud and edge/on-prem infrastructure, especially industrial or manufacturing contexts ⢠Familiarity with Snowflake, BigQuery, Redshift, or Databricks ⢠Experience with service mesh, zero-trust networking, or industrial/critical-infrastructure compliance frameworks such as SOC 2 or IEC 62443 ⢠History of building internal developer platforms or self-service infrastructure tooling ⢠Experience scaling infrastructure teams or setting technical direction at Staff level
⢠Significant opportunities for career development and advancement ⢠Comprehensive salary and equity package ⢠Innovative environment focused on transforming heavy industries through AI and automation ⢠Collaborative culture valuing innovation, discipline, and continuous improvement
Apply Nowđ August 18
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
đ August 5
Staff SRE leading BeyondTrustâs Password Safe identity-security platform across cloud and on-premises infrastructure. Driving GitOps, CI/CD, observability, resilience, and reliability strategy.
đ¨đŚ Canada â Remote
đ° Private Equity Round on 2021-05
â° Full Time
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đ July 16
Director of SRE leading global infrastructure and reliability teams for Blackpoint Cyber's cyber defense services. Focusing on cost efficiency, automation, and team leadership.
đ¨đŚ Canada â Remote
đľ CA$167k - CA$213k / year
đ° $190M Series C on 2023-06
â° Full Time
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đ July 10
Managed Services Reliability Engineer supporting Canadian customersâ AWS cloud infrastructure at OpsGuru. Leading incident response, troubleshooting, security, backup, and reliability operations.
đ¨đŚ Canada â Remote
đľ $140k / year
đ° Private Equity Round on 2019-01
â° Full Time
đ Senior
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đ July 9
Senior technical leader responsible for designing cloud infrastructure and advancing DevOps practices at Conga. Collaborating with Engineering, Product, Security, and Operations teams to enhance software delivery.