Site Reliability Engineer – Level 3

Vaga não está no LinkedIn

🕒 Julho 29

🇮🇳 Índia – Remoto

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 15%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Granicus

Granicus

501 - 1000 funcionários

Fundada em 1999

🏛️ Governo

☁️ SaaS

📋 Conformidade

Government • SaaS • Compliance

A Granicus é uma empresa de tecnologia focada no governo que fornece uma Government Experience Cloud e uma gama de serviços digitais para agências locais, estaduais, federais, de educação e de distritos especiais. Seus produtos incluem plataformas de engajamento e comunicação, nuvens de serviço e operações (para licenças, registros, solicitações de serviço/311), gerenciamento de reuniões e agendas, sites/CMS, ferramentas de conformidade e um Agente de Experiência do Governo com tecnologia AI para oferecer autoatendimento 24 horas por dia, 7 dias por semana. A Granicus ajuda organizações do setor público a modernizar a prestação de serviços, aumentar o engajamento dos cidadãos, automatizar fluxos de trabalho e melhorar a eficiência operacional.

Descrição

• Provide on-call production support, ensuring rapid triage, escalation handling, and service restoration. • Investigate production and customer issues, lead incident troubleshooting, and drive rapid RCA with clear follow-ups. • Use AIOps-assisted RCA, log clustering, timeline reconstruction, and incident summarization to accelerate diagnosis while validating AI recommendations before action. • Own and evolve the observability stack, with deep expertise in ELK/OpenSearch (Elasticsearch, Logstash, Kibana) for log ingestion, indexing, querying, visualization, and alerting. • Design and maintain observability across logs, metrics, and traces, ensuring actionable monitoring and high signal-to-noise alerting. • Build and enhance workflows for alerting, anomaly detection, and incident enrichment to reduce noise and improve accuracy. • Implement AIOps use cases such as dynamic baselining, alert correlation, event suppression, impact prediction, and automated context enrichment. • Develop automation, runbooks, and controlled self-healing mechanisms with appropriate safeguards, rollback plans, and auditability. • Implement AIOps remediation patterns that connect observability signals to runbooks, tickets, ChatOps actions, and human-approved recovery steps. • Drive improvements in system reliability, performance, scalability, and resilience through engineering-led initiatives. • Partner with engineering teams to improve deployment safety, operational readiness, and production stability. • Maintain high-quality runbooks, documentation, and knowledge bases to improve on-call effectiveness and knowledge sharing. • Support capacity planning, performance tuning, and SLO-based reliability practices. • Apply security, access control, and operational guardrails across systems and automation. • Own AIOps implementation from use-case definition through production rollout, including telemetry readiness, model/rule configuration, integration testing, operational validation, adoption tracking, and continuous tuning.

🎯 Requisitos

• 6+ years of experience in SRE, AIOps, or production engineering in large-scale, cloud environments. • Strong expertise in Linux/Unix, networking, distributed systems, and cloud platforms (AWS/Azure/GCP) . • Expert in ELK/OpenSearch, including: Log ingestion (Logstash / Beats) Elasticsearch index design, scaling, and tuning Advanced Kibana querying and debugging Dashboards, alerts, and observability patterns for production systems • Hands-on experience in logs, metrics, and tracing . • Ability to prepare telemetry for AIOps implementation, including tagging, normalization, correlation keys, service mapping, and high-quality event metadata. • Solid understanding of incident management, RCA, SLOs, and operational best practices . • Good understanding of AIOps: anomaly detection, alert correlation, and intelligent alerting . • Hands-on implementation experience with AIOps workflows, including event correlation, signal enrichment, noise suppression, automated incident summaries, and governed remediation. • Ability to implement AIOps integrations across observability tools, ITSM/ticketing systems, ChatOps, CMDB/runbook repositories, and automation platforms. • Experience measuring AIOps effectiveness through operational KPIs such as alert-noise reduction, faster MTTD/MTTR, RCA quality, automation adoption, repeat usage, and business impact linkage. • Experience with Infrastructure as Code tools such as Terraform, Ansible, or similar. • Preferred Certifications : AWS DevOps Engineer, AWS ML Specialty, Google Cloud DevOps Engineer, Azure DevOps Engineer, Kubernetes/CKA, or relevant AI/ML, AIOps, observability, or cloud automation certifications.

🏖️ Benefícios

• Employee Resource Groups to encourage diverse voices • Coffee with Mark sessions – Our employees get to interact with our CEO on very important and sometimes difficult issues ranging from mental health to work-life balance and current affairs. • Microsoft Teams communities focused on wellness, art, furbabies, family, parenting, and more. • Special guests from time to time to discuss issues that impact our employee population

Candidatar-se

Vagas Similares

🕒 Julho 28

Sezzle

201 - 500

💳 Fintech

👥 B2C

🛍️ Comércio Eletrônico

Senior Site Reliability Engineer at Sezzle resolving infrastructure challenges and enhancing reliability through scalable solutions. Seeking innovative and experienced candidates to drive technical excellence.

🇮🇳 Índia – Remoto

💵 $5.000 - $9.500 / mês

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

fal

51 - 200

🤖 Inteligência Artificial

🔌 API

☁️ SaaS

Machine Learning Engineer focusing on the reliability and security of generative media model APIs at fal. Working with cutting-edge models and infrastructure in a remote setting.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

BCD Travel

10.000+ funcionários

💼 Consultoria

📦 Logística

🏨 Hospitalidade

DevOps Engineer responsible for designing, building, and enhancing automation solutions using Microsoft Power Automate. Collaborating with global teams to improve workflows and operational efficiency.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 27

QuantumLoopAi

51 - 200

🏥 Saúde

🤖 Inteligência Artificial

☁️ SaaS

Senior Azure DevOps Engineer responsible for managing and optimising Azure cloud infrastructure. Join QuantumLoopAI, a healthtech company scaling its AI-platform solutions.

🇮🇳 Índia – Remoto

💰 $2.000.000 Convertible note em 2025-02

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 27

Coforgetech Ltd

-

🤖 Inteligência Artificial

🤝 B2B

☁️ SaaS

DevOps Engineer responsible for improving customer experience through deployments and integrations. Collaborate on technical support and back-end system integration for enhanced operational efficiency.

🗣️🇺🇸🇬🇧 Inglês obrigatório