Senior Site Reliability Engineer

🕒 August 24

🇼🇳 India – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đŸ‘» Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Akamai Technologies

Akamai Technologies

5001 - 10000 employees

🔒 Cybersecurity

💰 Post-IPO Equity on 2001-07

Cloud Computing ‱ Cybersecurity ‱ Content Delivery

Akamai Technologies is a leading cloud services provider that specializes in delivering security, cloud computing, and content delivery solutions. It offers a range of services such as API security, DDoS protection, and performance optimization for web applications, ensuring secure and reliable user experiences. With a robust global infrastructure, Akamai empowers businesses to streamline their digital presence while safeguarding against various cyber threats and enhancing application performance.

📋 Description

‱ Oversee, scale, and optimize next-generation dedicated AI hardware infrastructure ‱ Ensure uptime and reliability of AI hardware infrastructure offerings ‱ Enhance reliability, scalability, and performance across high-density hardware and software infrastructure in regional data centers ‱ Define KPIs, implement proactive monitoring, automate operations, and resolve urgent issues ‱ Develop and scale Python tooling and infrastructure-as-code utilities to eliminate operational toil and automate fleet-wide provisioning ‱ Integrate automated workflows across corporate ticketing systems for hardware and network break-fix incidents ‱ Use AI tools and LLM-based development approaches for technical execution, script creation, and system evaluation ‱ Improve availability, latency, and systemic health of private cloud and compute environments ‱ Design telemetry pipelines, Prometheus/Grafana dashboards, and AI-based anomaly detection for bare-metal and virtualized environments ‱ Participate in 24x7x365 on-call rotations and lead real-time incident management through PagerDuty and Slack workflows ‱ Partner with infrastructure vendors and coordinate on-site field technicians

🎯 Requirements

‱ 5+ years of relevant experience ‱ Bachelor's degree in Computer Science or related field ‱ Exceptional proficiency in Python and tooling/coding for scalable operational tools, API integrations, and automation frameworks ‱ Hands-on experience with Prometheus, Grafana, OpenTelemetry, and Loki ‱ Working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks ‱ Expertise designing new service rollouts, operational readiness criteria, telemetry baselines, and alerting thresholds ‱ Extensive experience building technical runbooks, leading complex incident response bridges, and conducting blameless post-mortems ‱ Ability to own ambiguous technical challenges, coordinate cross-functional teams, and drive production-grade solutions

đŸ–ïž Benefits

‱ Health and well-being benefits ‱ Financial benefits ‱ FlexBase workplace flexibility: work at home, in an office, or a combination of both

Apply Now

Similar Jobs

🕒 August 21

Simbian

11 - 50

đŸ€– Artificial Intelligence

🔒 Cybersecurity

Forward Deployment Engineer implementing Simbian’s AI SOC platform for enterprise and MSSP customers in India. Integrating SIEM, EDR, XDR, IAM, cloud tools, APIs, and webhooks.

Cloud

Cyber Security

🕒 August 13

Lingaro

1001 - 5000

đŸ’Œ Consulting

📣 Marketing

📩 Logistics

DevOps Engineer maintaining observability platforms and infrastructure for Lingaros Group in India. Building Grafana and Prometheus monitoring, CI pipelines, automated diagnostics, and failover systems.

Grafana

Prometheus

🕒 August 12

Provenir

201 - 500

☁ SaaS

💳 Fintech

đŸ€– Artificial Intelligence

Senior DevOps Engineer architecting Provenir’s AWS and Kubernetes platform for AI-powered decision intelligence. Driving GitOps, AIOps, CI/CD modernization, data infrastructure, observability, and cloud cost optimization.

Ansible

AWS

Grafana

Jenkins

Kafka

Kubernetes

Postgres

Python

SDLC

Terraform

🕒 August 5

Empower

10,000+ employees

💾 Finance

💳 Fintech

đŸ‘„ B2C

Senior DevOps Engineer automating AWS infrastructure, CI/CD, and security for Empower’s financial services software. Building internal tools and improving application reliability for engineering teams.

Angular

Ansible

Apache

AWS

Azure

Chef

Cloud

DNS

Docker

DynamoDB

EC2

Google Cloud Platform

Java

JavaScript

Jenkins

jQuery

Kubernetes

Linux

Maven

MySQL

NGINX

Node.js

Python

React

Ruby

Splunk

Terraform

Go

🕒 August 5

Signalmash

51 - 200

đŸ’Œ Consulting

📩 Logistics

đŸ„ Healthcare

DevOps Engineer owning Kubernetes, CI/CD, PostgreSQL, observability, and security for Signalmash’s cloud communications platform. Improving reliability, deployment speed, recovery, and infrastructure costs from India.

AWS

Azure

Cloud

Docker

Flux

Google Cloud Platform

Grafana

JavaScript

Kubernetes

Linux

Node.js

Postgres

Prometheus

Python

Shell Scripting