Senior Site Reliability Engineer, Technical Leader – Kubernetes Platform

🔥 11 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cisco

Cisco

10,000+ employees

Founded 1984

🔧 Hardware

🔐 Security

🏢 Enterprise

Hardware • Security • Enterprise

Cisco is a multinational technology company that provides networking hardware, software, and services to enterprises, service providers, and governments. It builds routers, switches, optical transceivers, programmable silicon, and edge computing platforms, and offers security, collaboration (Webex), observability, and AI-enabled software and support services to help organizations design, operate, and secure large-scale networks and data centers. Cisco also delivers professional services, training, and cloud-managed solutions to support digital transformation and AI-ready infrastructure.

📋 Description

• Design, build, and operate production-grade Kubernetes platforms in a regulated and non-regulated environments. • Improve system reliability through automation, thoughtful design, and continuous iteration • Define and drive SLOs, SLIs, and error budgets to guide reliability decisions • Build and evolve CI/CD pipelines that are secure, scalable, and easy to use • Implement robust observability (metrics, logs, traces) to make systems understandable and actionable • Reduce operational toil by automating repetitive processes and improving workflows • Partner with security and compliance teams to meet Compliance requirements without sacrificing developer velocity • Support audit processes, including documentation, controls implementation, and audit readiness • Participate in on-call rotations supporting customer requests and paging alerts • Participate in incident response, blameless postmortems, and continuous improvement efforts • Help shape a platform that engineers enjoy using

🎯 Requirements

• 10+ years of experience in SRE, DevOps, or infrastructure engineering • Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream) • Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals • Experience with Infrastructure as Code (Terraform preferred) • Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD) • Proficiency in scripting or programming (Python, Go) • Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK) • Working knowledge of compliance frameworks (e.g., PCI, ISO)

🏖️ Benefits

• Flexible work arrangements • Professional development opportunities

Apply Now

Similar Jobs

🔥 8 hours ago

3Pillar Global

1001 - 5000

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Senior DevOps Engineer building AI-native products at 3Pillar. Collaborating with global teams and leading automation and CI/CD efforts.

Ansible

AWS

Azure

Chef

Cloud

Docker

Kubernetes

Linux

Prometheus

Puppet

Python

SQL

Terraform

🔥 13 hours ago

Convene

201 - 500

🏢 Enterprise

📋 Compliance

🔐 Security

Support and Deployment Engineer at Azeus Systems deploying cloud applications and managing server configurations. Work closely with project teams to ensure efficient system implementations.

AWS

Azure

Cloud

🕒 2 days ago

NIVA Health

51 - 200

⚕️ Healthcare Insurance

🧘 Wellness

DevOps Engineer building cloud platform for AI-powered healthcare solutions at NIVA Health. Designing CI/CD pipelines, managing Kubernetes, ensuring scalable and reliable deployments.

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform

🕒 2 days ago

Granicus

501 - 1000

🏛️ Government

☁️ SaaS

📋 Compliance

Senior DevOps Engineer focused on cloud automation and operational reliability for Govtech solutions. Leading technical projects and mentoring engineers in complex application environments.

Cloud

Linux

🕒 3 days ago

Red Hat

10,000+ employees

🏢 Enterprise

Customer Site Reliability Engineer for Red Hat's OpenShift Managed Cloud, focusing on reliability and performance. Manage distributed systems and drive continuous improvement for customer satisfaction.

Ansible

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

Kubernetes

Linux

OpenShift

Prometheus

TCP/IP

Terraform

Go