Site Reliability Engineer

🕒 Agosto 31

🏄 California – Remoto

infoinfo

💵 $160.000 - $200.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 17%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of AXON Networks

AXON Networks

201 - 500 funcionários

Fundada em 2021

💼 Consultoria

📦 Logística

📣 Marketing

Consulting • Logistics • Marketing

A AXON Networks é uma empresa de tecnologia focada no desenvolvimento de soluções inovadoras em redes. Especializa-se em fornecer opções avançadas de conectividade e sistemas de gerenciamento de dados adaptados para diversas indústrias, aprimorando a comunicação e a eficiência operacional.

Descrição

• Improve the availability, performance, scalability and recoverability of AXON Networks cloud solutions • Combine software engineering with hands-on NOC operations to make the complete cloud-to-device service path observable, supportable and resilient at fleet scale • Establish practical SRE capabilities inside the NOC while partnering with Support, Operations, cloud and DevOps Engineering • Own reliability outcomes for assigned cloud services • Improve observability, capacity, resilience and recovery • Define and operationalize service-level indicators, service-level objectives and actionable alerting • Automate repetitive NOC work and create safe, testable mechanisms for diagnosis, recovery, device operations and routine production changes • Lead technically during incidents, drive evidence-based learning and ensure corrective actions are completed • Establish reliability baselines, SLIs, SLOs and error budgets for cloud services and device-management workflows • Trace failures across cloud APIs, microservices, Kubernetes, infrastructure, databases, messaging, networks, device-management protocols and devices • Identify fleet-wide and customer-specific failure patterns • Contribute operability requirements and production evidence during design and readiness reviews • Maintain NOC dashboards for service health, device reachability, provisioning, command and telemetry performance, firmware adoption and customer impact • Participate in the NOC production on-call rotation and serve as technical incident lead or senior troubleshooter • Diagnose complex application, infrastructure, Kubernetes, API, networking, database, messaging and CPE failures • Coordinate evidence gathering and technical escalation with customers, Engineering, firmware, DevOps and vendors • Lead or contribute to post-incident reviews and prioritize measurable corrective actions • Develop production-grade software, scripts and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance and recovery • Improve CI/CD and GitOps practices, including automated testing, release validation, progressive delivery and rollback readiness • Manage or contribute to infrastructure as code, configuration as code and reusable self-service patterns • Measure NOC toil and prioritize durable platform capabilities • Develop capacity models for service-provider growth, managed devices, telemetry, messaging, API demand and rollout events • Create and maintain runbooks, troubleshooting decision trees, service maps, dependency records and operational knowledge • Coach NOC and Support personnel on diagnosis, mitigation, evidence capture and escalation • Build self-service diagnostic views and tools for determining scope, affected customers, device cohorts, fault domain and next action • Share reliability insights with Engineering and Product and contribute to reliability and operational-readiness reviews

🎯 Requisitos

• 5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role • Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code • Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network and device-integration layers • Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments • Experience with infrastructure as code and delivery tooling such as Terraform, Helm, Git-based CI/CD and policy-as-code • Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis • Strong troubleshooting and debugging skills in Kubernetes platforms • Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring • Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems • Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices • Experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents • Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning • Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and systematic packet- or session-level troubleshooting • Clear communication, disciplined documentation and the ability to collaborate across NOC, cloud, DevOps, firmware and service-provider teams • Bachelor’s degree in computer science, engineering or equivalent practical experience • Preferred: experience supporting cloud-managed CPEs, broadband gateways, routers, ONTs, Wi-Fi/mesh systems or similar edge devices • Preferred: familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management • Preferred: experience supporting Apache Pulsar or Kafka, APIs and highly available databases used in device-management control planes • Preferred: understanding of GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless access technologies • Preferred: experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and safe rollback practices • Preferred: experience building auto-remediation, safe self-service operations or internal reliability platforms • Preferred: experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment • Must be able to work without visa sponsorship

Candidatar-se

Vagas Similares

🕒 Agosto 28

ComPsych

1001 - 5000

🏥 Saúde

💼 Consultoria

📦 Logística

Senior DevSecOps Engineer modernizing cloud infrastructure and CI/CD for ComPsych, a workplace mental-health and absence-management provider. Designing secure automation, observability, and deployment standards across application teams.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 28

Penn Interactive

201 - 500

🎲 Jogos de Azar

🎮 Jogos

🛍️ Comércio Eletrônico

Senior SRE operating Kubernetes and cloud infrastructure for PENN Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production environments.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 28

Summit Racing Equipment

11 - 50

🚘 Automotivo

🛒 Varejo

🛍️ Comércio Eletrônico

DevOps Engineer supporting Summit Racing Equipment’s software delivery and reliability systems. Managing monitoring, CI/CD deployments, SDLC processes, and infrastructure-development team integration.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 28

ComPsych

1001 - 5000

🏥 Saúde

💼 Consultoria

📦 Logística

Senior DevOps Engineer modernizing cloud infrastructure and CI/CD at ComPsych, a workplace mental health and absence management provider. Designing secure automation, observability, and deployment standards across application teams.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 28

Experian

10.000+ funcionários

💼 Consultoria

📣 Marketing

📦 Logística

Mid DevOps Engineer modernizing AWS infrastructure and CI/CD for Experian’s global data and technology business. Supporting Audigent integration, cloud migrations, security, and efficiency.

🗣️🇺🇸🇬🇧 Inglês obrigatório