
5001 - 10000 employees
🔒 Cybersecurity
💰 Post-IPO Equity on 2001-07
Cloud Computing • Cybersecurity • Content Delivery
Akamai Technologies is a leading cloud services provider that specializes in delivering security, cloud computing, and content delivery solutions. It offers a range of services such as API security, DDoS protection, and performance optimization for web applications, ensuring secure and reliable user experiences. With a robust global infrastructure, Akamai empowers businesses to streamline their digital presence while safeguarding against various cyber threats and enhancing application performance.
🕒 July 29
Improve your chances of getting an interview by checking your resume score before you apply.

5001 - 10000 employees
🔒 Cybersecurity
💰 Post-IPO Equity on 2001-07
Cloud Computing • Cybersecurity • Content Delivery
Akamai Technologies is a leading cloud services provider that specializes in delivering security, cloud computing, and content delivery solutions. It offers a range of services such as API security, DDoS protection, and performance optimization for web applications, ensuring secure and reliable user experiences. With a robust global infrastructure, Akamai empowers businesses to streamline their digital presence while safeguarding against various cyber threats and enhancing application performance.
• Collaborate with cross-functional teams to optimize the performance, availability, and reliability of Akamai’s core Mapping Service. • Define critical KPIs, advance our monitoring and alerting infrastructure, and architect automated operational responses. • Apply statistical analysis and cutting-edge machine learning to diagnose and solve the internet’s most complex content delivery challenges. • Co-design, manage, and track product SLIs/SLOs to proactively monitor, investigate, and analyze system performance and availability. • Apply advanced analytical skills and statistical insights to identify mapping bottlenecks, resolve reliability challenges, and engineer long-term solutions. • Build and deploy internal tools that automate proactive performance tracking and accelerate independent incident diagnosis. • Partner with product engineers to champion scalable, resilient, and highly supportable system architectures. • Provide data-driven insights to guide executive-level decision-making and identify high-impact technology investments. • Collaborate with internal engineering teams to swiftly troubleshoot, root-cause, and resolve complex customer escalations.
• Master’s or PhD in Computer Science or a highly analytical equivalent field. • 5+ years of experience in Site Reliability Engineering (SRE) or a related engineering role. • Deep mastery of Unix/Linux internals, computer networking protocols, and distributed system design. • Professional fluency operating within a command-line UNIX/Linux computing environment. • Strong background in statistical data analysis, SQL database querying, and data integrity troubleshooting. • Proven ability to transform complex datasets into actionable strategic roadmaps. • Practical knowledge of enterprise observability, logging, and alerting systems like Grafana. • Coding proficiency in a major backend or scripting language (e.g., Python).
• We support your health, well-being, finances, and life beyond work. See our benefits. • FlexBase adapts to your job's needs • It's not about telling employees where to work; it's about supporting employees to do their best work.
Apply Now🕒 July 29
Site Reliability Engineer 3 modernizing reliability engineering for Granicus with a focus on AIOps and automation. Improve service reliability and build scalable, resilient platforms for various workloads.
Ansible
AWS
Azure
Cloud
Distributed Systems
ElasticSearch
Google Cloud Platform
ITSM
Kubernetes
Linux
Logstash
Terraform
Unix
🕒 July 28
Senior Site Reliability Engineer at Sezzle resolving infrastructure challenges and enhancing reliability through scalable solutions. Seeking innovative and experienced candidates to drive technical excellence.
AWS
Distributed Systems
Grafana
Kubernetes
Microservices
MySQL
Postgres
Prometheus
RDBMS
SQL
Go
🕒 July 28
Engineering Manager at Moniepoint driving delivery and execution for the Site Reliability Engineering team. Collaborating with cross-functional teams to ensure high technical standards and timely feature delivery.
JavaScript
Node.js
Python
🕒 July 28
Machine Learning Engineer focusing on the reliability and security of generative media model APIs at fal. Working with cutting-edge models and infrastructure in a remote setting.
Distributed Systems
🕒 July 28
Customer Deployment Engineer at Harrison.ai ensuring best deployment experience for healthcare AI solutions. Collaborating with customers and partners for successful integrations and support.
🇮🇳 India – Remote
💰 Series B on 2021-12
⏰ Full Time
🟠 Senior
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
DNS
Linux
Python
TCP/IP