Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

🏄 California – Remote

infoinfo

💵 $160k - $200k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of AXON Networks

AXON Networks

201 - 500 employees

Founded 2021

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

AXON Networks is a technology company focused on developing innovative networking solutions. They specialize in providing advanced connectivity options and data management systems tailored for various industries, enhancing communication and operational efficiency.

📋 Description

• Improve the availability, performance, scalability and recoverability of AXON Networks cloud solutions • Combine software engineering with hands-on NOC operations to make the complete cloud-to-device service path observable, supportable and resilient at fleet scale • Establish practical SRE capabilities inside the NOC while partnering with Support, Operations, cloud and DevOps Engineering • Own reliability outcomes for assigned cloud services • Improve observability, capacity, resilience and recovery • Define and operationalize service-level indicators, service-level objectives and actionable alerting • Automate repetitive NOC work and create safe, testable mechanisms for diagnosis, recovery, device operations and routine production changes • Lead technically during incidents, drive evidence-based learning and ensure corrective actions are completed • Establish reliability baselines, SLIs, SLOs and error budgets for cloud services and device-management workflows • Trace failures across cloud APIs, microservices, Kubernetes, infrastructure, databases, messaging, networks, device-management protocols and devices • Identify fleet-wide and customer-specific failure patterns • Contribute operability requirements and production evidence during design and readiness reviews • Maintain NOC dashboards for service health, device reachability, provisioning, command and telemetry performance, firmware adoption and customer impact • Participate in the NOC production on-call rotation and serve as technical incident lead or senior troubleshooter • Diagnose complex application, infrastructure, Kubernetes, API, networking, database, messaging and CPE failures • Coordinate evidence gathering and technical escalation with customers, Engineering, firmware, DevOps and vendors • Lead or contribute to post-incident reviews and prioritize measurable corrective actions • Develop production-grade software, scripts and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance and recovery • Improve CI/CD and GitOps practices, including automated testing, release validation, progressive delivery and rollback readiness • Manage or contribute to infrastructure as code, configuration as code and reusable self-service patterns • Measure NOC toil and prioritize durable platform capabilities • Develop capacity models for service-provider growth, managed devices, telemetry, messaging, API demand and rollout events • Create and maintain runbooks, troubleshooting decision trees, service maps, dependency records and operational knowledge • Coach NOC and Support personnel on diagnosis, mitigation, evidence capture and escalation • Build self-service diagnostic views and tools for determining scope, affected customers, device cohorts, fault domain and next action • Share reliability insights with Engineering and Product and contribute to reliability and operational-readiness reviews

🎯 Requirements

• 5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role • Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code • Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network and device-integration layers • Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments • Experience with infrastructure as code and delivery tooling such as Terraform, Helm, Git-based CI/CD and policy-as-code • Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis • Strong troubleshooting and debugging skills in Kubernetes platforms • Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring • Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems • Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices • Experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents • Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning • Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and systematic packet- or session-level troubleshooting • Clear communication, disciplined documentation and the ability to collaborate across NOC, cloud, DevOps, firmware and service-provider teams • Bachelor’s degree in computer science, engineering or equivalent practical experience • Preferred: experience supporting cloud-managed CPEs, broadband gateways, routers, ONTs, Wi-Fi/mesh systems or similar edge devices • Preferred: familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management • Preferred: experience supporting Apache Pulsar or Kafka, APIs and highly available databases used in device-management control planes • Preferred: understanding of GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless access technologies • Preferred: experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and safe rollback practices • Preferred: experience building auto-remediation, safe self-service operations or internal reliability platforms • Preferred: experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment • Must be able to work without visa sponsorship

Apply Now

Similar Jobs

🕒 2 days ago

NetCov

201 - 500

🔒 Cybersecurity

🤝 B2B

💼 Consulting

AI Deployment Engineer deploying secure AI solutions across Hatz AI, Microsoft Co-Pilot, and Anthropic Claude. Supporting NetCov’s IT and cybersecurity services through data governance, troubleshooting, and customer implementations.

🇺🇸 United States – Remote

💵 $90k - $150k / year

💰 Private equity on 2022-11

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 2 days ago

NetCov

201 - 500

💼 Consulting

🏥 Healthcare

🔒 Cybersecurity

AI Deployment Engineer deploying secure AI solutions across Hatz AI, Microsoft Co-Pilot, and Anthropic Claude. Supporting NetCov’s customer IT and cybersecurity environments through governance, troubleshooting, and integration.

🇺🇸 United States – Remote

💵 $90k - $150k / year

💰 Private equity on 2022-11

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 2 days ago

Avalon Healthcare Solutions

51 - 200

🏥 Healthcare

☁️ SaaS

⚕️ Healthcare Insurance

Site Reliability Engineer securing and scaling AWS infrastructure for Avalon Healthcare Solutions’ diagnostic intelligence platform. Automating cloud operations, observability, and network security with Terraform and Palo Alto.

🕒 2 days ago

Concept Plus, LLC

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Mid-Level DevOps Engineer building secure CI/CD and cloud infrastructure for federal agencies. Supporting classified and unclassified DoD environments at Concept Plus.

🕒 2 days ago

ComPsych

1001 - 5000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior DevSecOps Engineer modernizing cloud infrastructure and CI/CD for ComPsych, a workplace mental-health and absence-management provider. Designing secure automation, observability, and deployment standards across application teams.