Site Reliability Engineer – Level 3

Job not on LinkedIn

🔥 10 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Granicus

Granicus

501 - 1000 employees

Founded 1999

🏛️ Government

☁️ SaaS

📋 Compliance

Government • SaaS • Compliance

Granicus is a company focused on transforming the way governments interact with their constituents through digital services and technology solutions. It provides the Government Experience Cloud to improve service delivery, community engagement, and operational efficiency across local, state, and federal governments. Granicus offers tools for agenda and meeting management, digital communication and engagement, public records management, and more, all designed to enhance customer experience and foster transparent and equitable interactions between governments and the people they serve.

📋 Description

• Provide on-call production support, ensuring rapid triage, escalation handling, and service restoration. • Investigate production and customer issues, lead incident troubleshooting, and drive rapid RCA with clear follow-ups. • Use AIOps-assisted RCA, log clustering, timeline reconstruction, and incident summarization to accelerate diagnosis while validating AI recommendations before action. • Own and evolve the observability stack, with deep expertise in ELK/OpenSearch (Elasticsearch, Logstash, Kibana) for log ingestion, indexing, querying, visualization, and alerting. • Design and maintain observability across logs, metrics, and traces, ensuring actionable monitoring and high signal-to-noise alerting. • Build and enhance workflows for alerting, anomaly detection, and incident enrichment to reduce noise and improve accuracy. • Implement AIOps use cases such as dynamic baselining, alert correlation, event suppression, impact prediction, and automated context enrichment. • Develop automation, runbooks, and controlled self-healing mechanisms with appropriate safeguards, rollback plans, and auditability. • Implement AIOps remediation patterns that connect observability signals to runbooks, tickets, ChatOps actions, and human-approved recovery steps. • Drive improvements in system reliability, performance, scalability, and resilience through engineering-led initiatives. • Partner with engineering teams to improve deployment safety, operational readiness, and production stability. • Maintain high-quality runbooks, documentation, and knowledge bases to improve on-call effectiveness and knowledge sharing. • Support capacity planning, performance tuning, and SLO-based reliability practices. • Apply security, access control, and operational guardrails across systems and automation. • Own AIOps implementation from use-case definition through production rollout, including telemetry readiness, model/rule configuration, integration testing, operational validation, adoption tracking, and continuous tuning.

🎯 Requirements

• 6+ years of experience in SRE, AIOps, or production engineering in large-scale, cloud environments. • Strong expertise in Linux/Unix, networking, distributed systems, and cloud platforms (AWS/Azure/GCP) . • Expert in ELK/OpenSearch, including: Log ingestion (Logstash / Beats) Elasticsearch index design, scaling, and tuning Advanced Kibana querying and debugging Dashboards, alerts, and observability patterns for production systems • Hands-on experience in logs, metrics, and tracing . • Ability to prepare telemetry for AIOps implementation, including tagging, normalization, correlation keys, service mapping, and high-quality event metadata. • Solid understanding of incident management, RCA, SLOs, and operational best practices . • Good understanding of AIOps: anomaly detection, alert correlation, and intelligent alerting . • Hands-on implementation experience with AIOps workflows, including event correlation, signal enrichment, noise suppression, automated incident summaries, and governed remediation. • Ability to implement AIOps integrations across observability tools, ITSM/ticketing systems, ChatOps, CMDB/runbook repositories, and automation platforms. • Experience measuring AIOps effectiveness through operational KPIs such as alert-noise reduction, faster MTTD/MTTR, RCA quality, automation adoption, repeat usage, and business impact linkage. • Experience with Infrastructure as Code tools such as Terraform, Ansible, or similar. • Preferred Certifications : AWS DevOps Engineer, AWS ML Specialty, Google Cloud DevOps Engineer, Azure DevOps Engineer, Kubernetes/CKA, or relevant AI/ML, AIOps, observability, or cloud automation certifications.

🏖️ Benefits

• Employee Resource Groups to encourage diverse voices • Coffee with Mark sessions – Our employees get to interact with our CEO on very important and sometimes difficult issues ranging from mental health to work-life balance and current affairs. • Microsoft Teams communities focused on wellness, art, furbabies, family, parenting, and more. • Special guests from time to time to discuss issues that impact our employee population

Apply Now

Similar Jobs

🔥 21 hours ago

Sezzle

201 - 500

💳 Fintech

👥 B2C

🛍️ eCommerce

Senior Site Reliability Engineer at Sezzle resolving infrastructure challenges and enhancing reliability through scalable solutions. Seeking innovative and experienced candidates to drive technical excellence.

AWS

Distributed Systems

Grafana

Kubernetes

Microservices

MySQL

Postgres

Prometheus

RDBMS

SQL

Go

🔥 23 hours ago

Moniepoint Inc. (Formerly TeamApt Inc.)

1001 - 5000

💳 Fintech

🏦 Banking

Engineering Manager at Moniepoint driving delivery and execution for the Site Reliability Engineering team. Collaborating with cross-functional teams to ensure high technical standards and timely feature delivery.

JavaScript

Node.js

Python

🕒 Yesterday

Ford Motor Company

10,000+ employees

📦 Logistics

💼 Consulting

📣 Marketing

DevSecOps Full Stack Engineer responsible for designing and delivering security products and platforms. Collaborating cross-functionally to automate security processes and embed best practices.

Cloud

🕒 Yesterday

fal

51 - 200

🤖 Artificial Intelligence

🔌 API

☁️ SaaS

Machine Learning Engineer focusing on the reliability and security of generative media model APIs at fal. Working with cutting-edge models and infrastructure in a remote setting.

Distributed Systems

🕒 Yesterday

harrison.ai

51 - 200

🏥 Healthcare

🤖 Artificial Intelligence

☁️ SaaS

Customer Deployment Engineer at Harrison.ai ensuring best deployment experience for healthcare AI solutions. Collaborating with customers and partners for successful integrations and support.

DNS

Linux

Python

TCP/IP