Staff Site Reliability Operations Engineer

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Calix

Calix

1001 - 5000 employees

Founded 2000

📡 Telecommunications

☁️ SaaS

🏢 Enterprise

💰 $50M Venture Round on 2009-08

Telecommunications • SaaS • Enterprise

Calix is a comprehensive solutions provider that focuses on enabling broadband service providers (BSPs) to simplify, innovate, and grow their businesses. Through its advanced broadband platform, Calix offers technologies like 10G-PON, 5G, Wi-Fi 7, and more to improve network operations, reduce downtime, and enhance the subscriber experience. The company provides managed services such as SmartLife and SmartHome, which help subscribers operate, secure, and enhance their connected lifestyles. Calix serves diverse provider types including telcos, cable operators, and municipal utilities, helping them deliver critical broadband connectivity and transform community access to digital services. With a focus on cloud technology, analytics, and transformation guidance, Calix empowers service providers to thrive in the digital age.

📋 Description

• Lead global platform reliability and observability strategy on GCP • Architect and troubleshoot complex networking infrastructure • Design and optimize observability platform using Grafana • Deploy machine learning models for anomaly detection • Drive architecture and security of GKE clusters • Ensure performance and scalability of databases (PostgreSQL, AlloyDB, BigQuery) • Automate incident response and triaging using AIOps • Mentor engineers on advanced debugging techniques

🎯 Requirements

• 8+ years in SRE, Production Engineering, or Distributed Systems • Deep technical knowledge across OSI layers (L1-L7) • Expert-level mastery of Google Kubernetes Engine (GKE) • Proven track record managing Apache Kafka pipelines and large-scale data environments • Experience deploying and managing Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo at scale • Advanced expertise utilizing HashiCorp Terraform for GCP architectures • High proficiency in Go and Python

🏖️ Benefits

• Health insurance • Bonuses

Apply Now

Similar Jobs

🔥 1 hour ago

Datavant

201 - 500

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Staff Site Reliability Engineer at Datavant focuses on cloud infrastructure design and security. Collaborating with teams to enhance reliability, scalability, and operability in healthcare data solutions.

Ansible

AWS

Azure

Cloud

DNS

Terraform

🔥 1 hour ago

Skylo

51 - 200

📡 Telecommunications

🤝 B2B

Staff Network Reliability Engineer managing RAN operations for satellite connectivity at Skylo. Overseeing RAN health and incident resolution in a production NTN environment.

🇺🇸 United States – Remote

💵 $150k - $162k / year

💰 $30M Venture Round - Skylo on 2025-02

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

Grafana

IoT

Kubernetes

Node.js

Prometheus

TypeScript

🔥 4 hours ago

CDW

10,000+ employees

💼 Consulting

🏥 Healthcare

📚 Education

Principal Consulting Engineer optimizing private cloud infrastructure built on VCF 9.0 for CDW. Driving standardization and automation across compute, storage, and networking management layers.

Cloud

🔥 6 hours ago

RTX

10,000+ employees

🚀 Aerospace

🎖️ Defense

🏭 Manufacturing

Site Reliability Engineer automating and improving reliability at Collins Aerospace. Collaborating with teams to tackle complex technical problems in flight operations.

🇺🇸 United States – Remote

💵 $107.5k - $204.5k / year

💰 $200k Grant - RTX on 2024-11

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

Docker

Kubernetes

Linux

SaltStack

Terraform

Unix

🔥 14 hours ago

Ad Hoc LLC

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Staff DevOps Engineer at Ad Hoc shaping the long-term technical strategy and mentoring team members. Leading critical projects and ensuring compliance in delivering software for Veterans Affairs.

AWS

Cloud

HAProxy

Kubernetes

Terraform