Staff Site Reliability Operations Engineer

🕒 2 days ago

🌐 United States, Canada – Remote

info

💵 $136k - $231k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Calix

Calix

1001 - 5000 employees

Founded 2000

📡 Telecommunications

☁️ SaaS

🏢 Enterprise

💰 $50M Venture Round on 2009-08

Telecommunications • SaaS • Enterprise

Calix is a comprehensive solutions provider that focuses on enabling broadband service providers (BSPs) to simplify, innovate, and grow their businesses. Through its advanced broadband platform, Calix offers technologies like 10G-PON, 5G, Wi-Fi 7, and more to improve network operations, reduce downtime, and enhance the subscriber experience. The company provides managed services such as SmartLife and SmartHome, which help subscribers operate, secure, and enhance their connected lifestyles. Calix serves diverse provider types including telcos, cable operators, and municipal utilities, helping them deliver critical broadband connectivity and transform community access to digital services. With a focus on cloud technology, analytics, and transformation guidance, Calix empowers service providers to thrive in the digital age.

📋 Description

• Architect, optimize, and troubleshoot complex Layer 1–Layer 7 networking infrastructure • Design, scale, and optimize the Grafana Labs observability stack, including Grafana, Mimir, Loki, Tempo, and Beyla • Deploy machine-learning models and automated anomaly detection to reduce alert fatigue and predict bottlenecks • Architect, scale, secure, and network production Google Kubernetes Engine clusters • Tune and maintain high-throughput Apache Kafka clusters • Ensure performance, scalability, and disaster-recovery readiness across PostgreSQL, AlloyDB, and BigQuery • Integrate AIOps insights with Grafana workflows to automate triage, root-cause analysis, and remediation • Champion the technical roadmap for distributed infrastructure engineering and GCP cloud-native observability standards • Mentor senior and junior engineers on advanced debugging, distributed systems, and intelligent operations

🎯 Requirements

• 8+ years of experience in SRE, Production Engineering, or Distributed Systems infrastructure roles • Proven success and autonomy in a 100% remote engineering environment • Deep technical knowledge across OSI Layers 1–7 • Physical/fiber infrastructure awareness, switching, BGP, and OSPF • TCP congestion control, UDP, and QUIC tuning • Session management, TLS termination, DNS architecture, HTTP/3, and gRPC • Expert mastery of GKE internals, custom controllers, multi-cluster networking, and GitOps workflows • Experience managing high-throughput Apache Kafka pipelines • Experience managing PostgreSQL, AlloyDB, and BigQuery at scale • Hands-on experience with Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo • Experience applying AI/ML to time-series anomaly detection, log clustering, and correlation • Advanced production-scale HashiCorp Terraform expertise for multi-region GCP architectures • High proficiency in Go and Python • Exceptional written and verbal communication skills • Deep knowledge of Google Cloud architecture, Cloud SDN, Cloud Armor, Interconnect, IAM, and cost optimization • Understanding of Linux internals, eBPF-based monitoring, kernel-level networking, Wireshark, and tcpdump

🏖️ Benefits

• Bonus eligibility as part of the total compensation package • Benefits package referenced by the employer

Apply Now

Similar Jobs

🕒 3 days ago

Ryan

1001 - 5000

💼 Consulting

🛡️ Insurance

💸 Finance

DevOps and cloud platform manager leading CI/CD, developer portals, and Azure/AWS automation. Managing engineers and advancing reliability, observability, security, and developer productivity for Ryan’s tax services.

🕒 3 days ago

Ryan

1001 - 5000

💼 Consulting

🛡️ Insurance

💸 Finance

DevOps and cloud platform manager leading CI/CD, cloud automation, and developer platforms. Advancing reliable, secure infrastructure for Ryan, a global tax services firm.

🕒 3 days ago

insightsoftware

1001 - 5000

☁️ SaaS

💸 Finance

🏢 Enterprise

Director leading global DevOps and CloudOps teams for insightsoftware’s financial reporting and analytics solutions. Managing cloud infrastructure, automation, reliability, security audits, and operational delivery.

🇺🇸 United States – Remote

💵 $184k - $231k / year

💰 $798.6M Private Equity Round - insightsoftware on 2021-07

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 3 days ago

3M Consultancy

1 - 10

💼 Consulting

📣 Marketing

🎯 Recruiter

DevOps Engineer automating secure IRS system integration and deployments for 3M Consultancy, a government/military consultancy. Managing cloud infrastructure, containers, CI/CD, and compliance-focused systems.

🇺🇸 United States – Remote

💵 $100k - $150k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 3 days ago

jrni

51 - 200

☁️ SaaS

🤝 B2B

🏢 Enterprise

Director leading DevOps, security, and IT for jrni’s enterprise appointment and queue-management platform. Owning AWS/GCP reliability, compliance, incident response, AI governance, and corporate IT across the US remotely.

🇺🇸 United States – Remote

💵 $165k - $185k / year

💰 $6M Series C - JRNI on 2019-08

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)