Site Reliability Engineer

🔥 1 minute ago

🇮🇳 India – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Granicus

Granicus

501 - 1000 employees

Founded 1999

🏛️ Government

☁️ SaaS

📋 Compliance

Government • SaaS • Compliance

Granicus is a company focused on transforming the way governments interact with their constituents through digital services and technology solutions. It provides the Government Experience Cloud to improve service delivery, community engagement, and operational efficiency across local, state, and federal governments. Granicus offers tools for agenda and meeting management, digital communication and engagement, public records management, and more, all designed to enhance customer experience and foster transparent and equitable interactions between governments and the people they serve.

📋 Description

• Provide on-call production support, rapid triage, escalation handling, and service restoration • Investigate production and customer issues, lead incident troubleshooting, and drive root-cause analysis • Use AIOps-assisted RCA, log clustering, timeline reconstruction, and incident summarization • Own and evolve the ELK/OpenSearch observability stack • Design and maintain observability across logs, metrics, and traces • Build workflows for alerting, anomaly detection, and incident enrichment • Implement dynamic baselining, alert correlation, event suppression, impact prediction, and automated context enrichment • Develop automation, runbooks, and controlled self-healing mechanisms with safeguards, rollback plans, and auditability • Connect observability signals to runbooks, tickets, ChatOps actions, and human-approved recovery steps • Improve system reliability, performance, scalability, and resilience • Partner with engineering teams on deployment safety, operational readiness, and production stability • Maintain runbooks, documentation, and knowledge bases • Support capacity planning, performance tuning, and SLO-based reliability practices • Apply security, access control, and operational guardrails • Own AIOps implementation from use-case definition through production rollout, validation, adoption tracking, and continuous tuning

🎯 Requirements

• 4+ years of experience in SRE, AIOps, or production engineering in large-scale, cloud environments • Strong expertise in Linux/Unix, networking, distributed systems, and AWS/Azure/GCP • Expert knowledge of ELK/OpenSearch, including Logstash/Beats, Elasticsearch index design, scaling and tuning, Kibana querying and debugging, dashboards, and alerts • Hands-on experience with logs, metrics, and tracing • Ability to prepare telemetry for AIOps implementation, including tagging, normalization, correlation keys, service mapping, and event metadata • Understanding of incident management, RCA, SLOs, and operational best practices • Hands-on AIOps implementation experience, including anomaly detection, event correlation, signal enrichment, noise suppression, incident summaries, and governed remediation • Ability to implement integrations across observability tools, ITSM/ticketing systems, ChatOps, CMDB/runbook repositories, and automation platforms • Experience measuring AIOps effectiveness using operational KPIs • Experience with Infrastructure as Code tools such as Terraform, Ansible, or similar • Preferred certifications include AWS DevOps Engineer, AWS ML Specialty, Google Cloud DevOps Engineer, Azure DevOps Engineer, Kubernetes/CKA, or relevant AI/ML, AIOps, observability, or cloud automation certifications • Shift work – EMEA shifts

🏖️ Benefits

• Remote-first company • Remote work arrangement • Employee Resource Groups • Coffee with Mark sessions • Microsoft Teams communities focused on wellness, art, furbabies, family, and parenting • Special guest discussions on employee-impacting issues

Apply Now

Similar Jobs

🔥 10 hours ago

4Pharma Ltd

11 - 50

💼 Consulting

🍽️ Food & Beverage

📦 Logistics

Senior DevOps Engineer building AWS, Kubernetes and CI/CD infrastructure for BC Platforms’ global healthcare data and analytics platform. Driving reliability, security and automation.

AWS

Azure

Cloud

Grafana

Kubernetes

Linux

Prometheus

🔥 20 hours ago

MFSG

11 - 50

🏭 Manufacturing

🔧 Hardware

🚗 Transport

Site Reliability Engineer automating reliable, compliant digital banking platforms for MFSG Technologies. Managing CI/CD, observability, incident response, and resilient production deployments.

Ansible

AWS

Azure

Cloud

Docker

Kubernetes

Python

Terraform

🔥 22 hours ago

Astreya

1001 - 5000

💼 Consulting

📦 Logistics

📣 Marketing

Lead DevOps engineer running Astreya's GCP infrastructure, deployments, security, and platform reliability. Mentoring engineers and supporting customer cloud environments.

AWS

Azure

Cloud

Distributed Systems

DNS

Google Cloud Platform

ITSM

Kubernetes

Python

ServiceNow

SQL

Terraform

Go

🕒 Yesterday

iCert Global

51 - 200

💼 Consulting

📣 Marketing

📚 Education

Lead SRE managing Azure infrastructure, AKS, observability, and major incidents for Icertis’s AI-powered contract intelligence platform. Driving automation, reliability, and cloud-native operations.

AWS

Azure

Cloud

Distributed Systems

Docker

Kubernetes

Python

ServiceNow

Terraform

🕒 2 days ago

Miratech

501 - 1000

🤝 B2B

💼 Consulting

☁️ SaaS

Senior Observability/DevOps Engineer migrating enterprise Datadog environments for Miratech, a global IT services and consulting company. Automating resilient observability platforms with Datadog, Terraform, AWS, and Kubernetes.

AWS

Kubernetes

ServiceNow

Terraform