Senior Site Reliability Engineer

🕒 July 29

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Akamai Technologies

Akamai Technologies

5001 - 10000 employees

🔒 Cybersecurity

💰 Post-IPO Equity on 2001-07

Cloud Computing • Cybersecurity • Content Delivery

Akamai Technologies is a leading cloud services provider that specializes in delivering security, cloud computing, and content delivery solutions. It offers a range of services such as API security, DDoS protection, and performance optimization for web applications, ensuring secure and reliable user experiences. With a robust global infrastructure, Akamai empowers businesses to streamline their digital presence while safeguarding against various cyber threats and enhancing application performance.

📋 Description

• Collaborate with cross-functional teams to optimize the performance, availability, and reliability of Akamai’s core Mapping Service. • Define critical KPIs, advance our monitoring and alerting infrastructure, and architect automated operational responses. • Apply statistical analysis and cutting-edge machine learning to diagnose and solve the internet’s most complex content delivery challenges. • Co-design, manage, and track product SLIs/SLOs to proactively monitor, investigate, and analyze system performance and availability. • Apply advanced analytical skills and statistical insights to identify mapping bottlenecks, resolve reliability challenges, and engineer long-term solutions. • Build and deploy internal tools that automate proactive performance tracking and accelerate independent incident diagnosis. • Partner with product engineers to champion scalable, resilient, and highly supportable system architectures. • Provide data-driven insights to guide executive-level decision-making and identify high-impact technology investments. • Collaborate with internal engineering teams to swiftly troubleshoot, root-cause, and resolve complex customer escalations.

🎯 Requirements

• Master’s or PhD in Computer Science or a highly analytical equivalent field. • 5+ years of experience in Site Reliability Engineering (SRE) or a related engineering role. • Deep mastery of Unix/Linux internals, computer networking protocols, and distributed system design. • Professional fluency operating within a command-line UNIX/Linux computing environment. • Strong background in statistical data analysis, SQL database querying, and data integrity troubleshooting. • Proven ability to transform complex datasets into actionable strategic roadmaps. • Practical knowledge of enterprise observability, logging, and alerting systems like Grafana. • Coding proficiency in a major backend or scripting language (e.g., Python).

🏖️ Benefits

• We support your health, well-being, finances, and life beyond work. See our benefits. • FlexBase adapts to your job's needs • It's not about telling employees where to work; it's about supporting employees to do their best work.

Apply Now

Similar Jobs

🕒 July 29

Granicus

501 - 1000

🏛️ Government

☁️ SaaS

📋 Compliance

Site Reliability Engineer 3 modernizing reliability engineering for Granicus with a focus on AIOps and automation. Improve service reliability and build scalable, resilient platforms for various workloads.

Ansible

AWS

Azure

Cloud

Distributed Systems

ElasticSearch

Google Cloud Platform

ITSM

Kubernetes

Linux

Logstash

Terraform

Unix

🕒 July 28

Sezzle

201 - 500

💳 Fintech

👥 B2C

🛍️ eCommerce

Senior Site Reliability Engineer at Sezzle resolving infrastructure challenges and enhancing reliability through scalable solutions. Seeking innovative and experienced candidates to drive technical excellence.

AWS

Distributed Systems

Grafana

Kubernetes

Microservices

MySQL

Postgres

Prometheus

RDBMS

SQL

Go

🕒 July 28

Moniepoint Inc. (Formerly TeamApt Inc.)

1001 - 5000

💳 Fintech

🏦 Banking

Engineering Manager at Moniepoint driving delivery and execution for the Site Reliability Engineering team. Collaborating with cross-functional teams to ensure high technical standards and timely feature delivery.

JavaScript

Node.js

Python

🕒 July 28

fal

51 - 200

🤖 Artificial Intelligence

🔌 API

☁️ SaaS

Machine Learning Engineer focusing on the reliability and security of generative media model APIs at fal. Working with cutting-edge models and infrastructure in a remote setting.

Distributed Systems

🕒 July 28

harrison.ai

51 - 200

🏥 Healthcare

🤖 Artificial Intelligence

☁️ SaaS

Customer Deployment Engineer at Harrison.ai ensuring best deployment experience for healthcare AI solutions. Collaborating with customers and partners for successful integrations and support.

DNS

Linux

Python

TCP/IP