Senior Site Reliability Engineer, Technical Leader – Kubernetes Platform

🕒 July 24

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 36%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cisco

Cisco

10,000+ employees

Founded 1984

🔧 Hardware

🔐 Security

🏢 Enterprise

Hardware • Security • Enterprise

Cisco is a multinational technology company that provides networking hardware, software, and services to enterprises, service providers, and governments. It builds routers, switches, optical transceivers, programmable silicon, and edge computing platforms, and offers security, collaboration (Webex), observability, and AI-enabled software and support services to help organizations design, operate, and secure large-scale networks and data centers. Cisco also delivers professional services, training, and cloud-managed solutions to support digital transformation and AI-ready infrastructure.

📋 Description

• Design, build, and operate production-grade Kubernetes platforms in regulated and non-regulated environments • Improve system reliability through automation, thoughtful design, and continuous iteration • Define and drive SLOs, SLIs, and error budgets to guide reliability decisions • Build and evolve secure, scalable, easy-to-use CI/CD pipelines • Implement observability using metrics, logs, and traces • Reduce operational toil by automating repetitive processes and improving workflows • Partner with security and compliance teams to meet compliance requirements without sacrificing developer velocity • Support audit processes, including documentation, controls implementation, and audit readiness • Participate in on-call rotations supporting customer requests and paging alerts • Participate in incident response, blameless postmortems, and continuous improvement efforts • Help shape a platform that engineers enjoy using

🎯 Requirements

• 10+ years of experience in SRE, DevOps, or infrastructure engineering • Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream) • Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals • Experience with Infrastructure as Code (Terraform preferred) • Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD) • Proficiency in scripting or programming (Python, Go) • Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK) • Working knowledge of compliance frameworks (e.g., PCI, ISO) • Demonstrated ability to influence technical direction and drive cross-functional initiatives across engineering, security, and operations teams • Experience mentoring engineers and providing technical leadership in SRE, platform engineering, or infrastructure teams

🏖️ Benefits

• Remote work arrangement • On-call rotation participation • Opportunities to grow and build solutions at global scale

Apply Now

Similar Jobs

🕒 July 21

Granicus

501 - 1000

🏛️ Government

☁️ SaaS

📋 Compliance

Senior DevOps Engineer focused on cloud automation and operational reliability for Govtech solutions. Leading technical projects and mentoring engineers in complex application environments.

Cloud

Linux

🕒 July 20

Five9

1001 - 5000

☁️ SaaS

🤖 Artificial Intelligence

📡 Telecommunications

Network Engineer maintaining scalable CI/CD pipelines for a cloud contact center software. Driving automation and network design in 24/7 SaaS environments.

Ansible

Cloud

Firewalls

Jenkins

Kubernetes

Linux

Python

Terraform

VoIP

🕒 July 18

BETSOL

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Cloud Engineer developing and operating cloud portal workloads across Azure and GCP using AI-first practices. Collaborating on security and automation in a global enterprise environment.

Ansible

AWS

Azure

Cloud

Google Cloud Platform

JavaScript

Jenkins

Kubernetes

Python

Terraform

TypeScript

🕒 July 16

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Site Reliability Engineer ensuring performance and reliability of Akamai's media and web delivery platform. Leading investigations into complex reliability and performance across global systems.

Distributed Systems

DNS

Linux

Python

SQL

TCP/IP

Unix

Go

🕒 July 13

MRSOOL | مرسول

201 - 500

🍽️ Food & Beverage

✈️ Travel

💼 Consulting

Site Reliability Engineer II for Mrsool, enhancing infrastructure and supporting development teams in a dynamic environment. Ensuring reliability for a leading delivery platform in the MENA region.

Ansible

AWS

Azure

Chef

Cloud

Distributed Systems

Docker

Google Cloud Platform

Grafana

Java

Kubernetes

Prometheus

Puppet

Python

Ruby

Terraform

Go