Site Reliability Engineer 3

Job not on LinkedIn

🕒 May 16

🇮🇳 India – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 45%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Granicus

Granicus

501 - 1000 employees

Founded 1999

🏛️ Government

☁️ SaaS

📋 Compliance

Government • SaaS • Compliance

Granicus is a company focused on transforming the way governments interact with their constituents through digital services and technology solutions. It provides the Government Experience Cloud to improve service delivery, community engagement, and operational efficiency across local, state, and federal governments. Granicus offers tools for agenda and meeting management, digital communication and engagement, public records management, and more, all designed to enhance customer experience and foster transparent and equitable interactions between governments and the people they serve.

📋 Description

• Provide production support according to the team on-call roster • Work on customer and internal engineering/implementation team tickets • Work on SRE backlog items • Monitor the health and performance of services, systems, and infrastructure • Respond promptly to alerts and incidents to maintain high availability • Develop and maintain automation scripts and tools • Troubleshoot and resolve incidents, perform root cause analysis, and implement long-term fixes • Design and implement system improvements for reliability, scalability, and performance • Collaborate with software engineers on application requirements, design, architecture, deployment, and releases • Create and maintain process, procedure, and troubleshooting documentation • Assist with capacity planning • Implement and follow security best practices to protect systems and data • Lead efforts to build and maintain robust infrastructure and guide implementation of site reliability best practices

🎯 Requirements

• Good understanding of Linux/Unix systems, networking, and cloud services such as AWS, Azure, or Google Cloud • Experience with Python, Bash, or Ruby • Bachelor’s or master’s degree in computer science, Information Technology, or a related field, or equivalent practical experience • 5+ years of experience in site reliability engineering, system administration, or a similar role • Proven track record managing large-scale, high-availability systems • Familiarity with AI/ML operations, including model lifecycle management, vector databases, and inference performance tuning • Expertise in Linux/Unix systems, networking, and cloud services • Proficiency in Python, Bash, Ruby, Go, Java, or C++ • Advanced knowledge of Elastic, Prometheus, Grafana, Splunk, Ansible, Chef, Puppet, and CI/CD pipelines • Strong analytical and problem-solving skills • Excellent verbal and written communication skills • Ability to lead and mentor a team, drive projects to completion, and manage cross-functional initiatives • Relevant certifications such as AWS Certified DevOps Engineer, AWS Certified Machine Learning – Specialty, or Google Cloud Professional DevOps Engineer are a plus

🏖️ Benefits

• Remote work / remote-first company • Employee Resource Groups • Coffee with Mark sessions with the CEO • Microsoft Teams communities focused on wellness, art, furbabies, family, and parenting • Special guest sessions addressing issues impacting employees

Apply Now

Similar Jobs

🕒 May 16

Proofpoint

1001 - 5000

🔒 Cybersecurity

🏢 Enterprise

🔐 Security

Site Reliability Engineer at Proofpoint managing and operating scalable distributed systems. Focused on Kubernetes infrastructure, CI/CD, and incident response processes across regions.

AWS

Cloud

Distributed Systems

Grafana

Kubernetes

Terraform

🕒 May 13

Shuru

51 - 200

🤖 Artificial Intelligence

🤝 B2B

🏢 Enterprise

Senior DevOps Engineer at Shuru Technologies enhancing cloud platform infrastructure. Collaborating with teams for scalable solutions and operational readiness in a remote-first environment.

AWS

Azure

Cloud

Google Cloud Platform

Kubernetes

Oracle

Postgres

Redis

SQL

Terraform

🕒 May 12

Volvo Cars

10,000+ employees

🏭 Manufacturing

🚗 Transport

🚘 Automotive

Salesforce Release Engineer driving digital innovation at Volvo Cars. Managing Salesforce release lifecycle across global teams and developing cutting-edge technology solutions for the automotive industry.

🕒 April 29

Tookitaki

51 - 200

🤖 Artificial Intelligence

Site Reliability Engineer maintaining and scaling infrastructure for fintech solutions at Tookitaki. Collaborating with engineering and DevOps teams for high availability and performance.

Ansible

AWS

Cloud

Docker

Google Cloud Platform

Grafana

Kubernetes

Linux

MariaDB

Prometheus

Python

SQL

Terraform

🕒 April 26

Avaya

5001 - 10000

💼 Consulting

📣 Marketing

📦 Logistics

Tooling Expert at Avaya serving as a technical liaison in Cloud Operations. Focusing on complex debugging, troubleshooting, and maintaining deployment tooling for CI/CD pipelines.

Cloud

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Python

Terraform