Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇨🇦 Canada – Remote

💵 $80 - $110 / hour

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of AXON Networks

AXON Networks

201 - 500 employees

Founded 2021

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

AXON Networks is a technology company focused on developing innovative networking solutions. They specialize in providing advanced connectivity options and data management systems tailored for various industries, enhancing communication and operational efficiency.

📋 Description

• Improve the availability, performance, scalability and recoverability of AXON Networks cloud solutions • Combine software engineering with hands-on NOC operations to make the cloud-to-device service path observable, supportable and resilient at fleet scale • Establish practical SRE capabilities inside the NOC while partnering with Support, Operations, cloud and DevOps Engineering • Participate in a sustainable on-call rotation and improve diagnosis of customer-impacting issues • Own reliability outcomes for assigned cloud services • Improve observability, capacity, resilience and recovery • Define and operationalize SLIs, SLOs and actionable alerting • Automate repetitive NOC work and create safe, testable mechanisms for diagnosis, recovery, device operations and routine production changes • Lead technically during incidents, drive evidence-based learning and complete corrective actions • Establish reliability baselines, SLIs, SLOs and error budgets for cloud services and device-management workflows • Trace failures across cloud APIs, microservices, Kubernetes, infrastructure, databases, messaging, networks, device-management protocols and devices • Identify fleet-wide and customer-specific failure patterns • Contribute operability requirements and production evidence during design and readiness reviews • Maintain NOC dashboards for service health, device reachability, provisioning, commands, telemetry, firmware adoption and customer impact • Participate in NOC production on-call rotation as technical incident lead or senior troubleshooter • Diagnose complex failures across applications, cloud infrastructure, Kubernetes, APIs, networking, DNS/TLS, databases, messaging, device-management sessions and CPE behavior • Coordinate evidence gathering and technical escalation with customers, Engineering, firmware, DevOps and vendors • Lead or contribute to post-incident reviews and convert recurring failures into measurable corrective actions • Develop production-grade software, scripts and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance and recovery • Improve CI/CD and GitOps practices, including automated testing, release validation, progressive delivery and rollback readiness • Manage or contribute to infrastructure as code, configuration as code and reusable self-service patterns • Measure NOC toil and prioritize durable platform capabilities with Automation & Tools Engineers • Develop capacity models for service-provider growth, managed-device populations, telemetry, messaging, APIs and rollout events • Create and maintain runbooks, troubleshooting decision trees, service maps, dependency records, known-error guidance and operational knowledge • Coach NOC and Support personnel on diagnosis, mitigation, evidence capture and escalation • Build self-service diagnostic views and tools for determining scope, affected customers, device cohorts, fault domain and next action • Share reliability insights with Engineering and Product and contribute to reliability and operational-readiness reviews

🎯 Requirements

• 5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role • Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code • Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting application, infrastructure, network and device-integration layers • Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments • Experience with Terraform, Helm, Git-based CI/CD and policy-as-code • Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis • Strong troubleshooting and debugging skills in Kubernetes platforms • Experience with observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring • Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems • Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices • Experience participating in an on-call rotation and responding to high-severity, customer-impacting production incidents • Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning • Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and packet- or session-level troubleshooting • Clear communication, disciplined documentation and collaboration across NOC, cloud, DevOps, firmware and service-provider teams • Bachelor’s degree in computer science, engineering or equivalent practical experience • Preferred: experience supporting cloud-managed CPEs in a service-provider environment • Preferred: familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management • Preferred: experience with Apache Pulsar or Kafka, APIs and highly available databases used in device-management control planes • Preferred: understanding of GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless access technologies • Preferred: experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and safe rollback practices • Preferred: experience building auto-remediation, safe self-service operations or internal reliability platforms • Preferred: experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment

🏖️ Benefits

• Equal opportunity recruitment process • Inclusive and diverse working environment

Apply Now

Similar Jobs

🕒 August 21

Workiy Inc.

11 - 50

💼 Consulting

📣 Marketing

🛍️ eCommerce

Senior DevOps Engineer owning Salesforce CI/CD architecture, release governance, sandbox strategies, and automated deployments. Introducing AI-assisted tools for testing, code review, and deployment risk detection.

🕒 August 7

PLATO

201 - 500

💼 Consulting

🤝 B2B

🏢 Enterprise

Senior DevOps Engineer building Azure DevOps pipelines, Terraform infrastructure, and deployment automation. Supporting PLATO, a Canadian Indigenous-owned software testing and technology services company, across product and data teams.

🕒 June 27

Workiy Inc.

11 - 50

💼 Consulting

📣 Marketing

🛍️ eCommerce

DevOps Engineer automating CI/CD pipelines and managing cloud infrastructure for major client projects. Join Workiy's remote team driving business success through technology expertise.

🕒 June 18

Workiy Inc.

11 - 50

💼 Consulting

📣 Marketing

🛍️ eCommerce

React/Python/AWS Engineer at Workiy designing cloud-native applications and integrating AI capabilities. Collaborating with stakeholders to drive innovation and best practices in software development.

🕒 May 20

High 5 Games

51 - 200

🎮 Gaming

🎲 Gambling

🤝 B2B

DevOps Engineer designing and optimizing cloud infrastructure for AI models on GCP. Collaborating with data scientists and engineers to maintain reliability and performance.