Senior Principal Site Reliability Engineer

🕒 March 27

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Akamai Technologies

Akamai Technologies

5001 - 10000 employees

🔒 Cybersecurity

🏱 Enterprise

đŸ“± Media

Cybersecurity ‱ Enterprise ‱ Media

Akamai Technologies is a global edge platform and cloud services company that delivers content delivery, edge computing, and security solutions. The company operates one of the world’s largest distributed networks to accelerate and protect web, media, and application traffic, offering products for content delivery, DDoS protection, API and app security, bot management, edge compute (serverless/edge functions), and AI inference at the edge. Akamai also provides enterprise-focused security services (zero trust, identity and access management, secure internet access) and cloud/AI infrastructure tools, and has recently expanded capabilities through acquisitions (for example LayerX) to add browser-based AI usage control.

📋 Description

‱ Defining the reliability architecture for Akamai's AI compute and platform services, including SLO frameworks, fault tolerance patterns, and capacity planning models ‱ Hands-on building of automation and tooling that reduces operational toil and scales the SRE team's impact ‱ Designing observability strategy by leveraging Akamai's existing platform to build the telemetry, dashboards, alerts, and GPU-specific monitoring needed for AI workloads ‱ Architecting deployment safety practices including progressive rollouts, canary analysis, rollback automation, and change safety processes ‱ Influencing product engineering architecture and design decisions, embedding reliability into the development lifecycle at the system level ‱ Mentoring and elevating other SREs through design reviews, code reviews, and hands-on problem-solving, setting the technical bar for the team

🎯 Requirements

‱ Have extensive experience in SRE, platform engineering, and/or infrastructure engineering, with demonstrated impact at a principal or staff level. ‱ Demonstrate extensive Kubernetes expertise, managing autoscaling, resource scheduling, and container orchestration for handling compute-intensive workloads effectively. ‱ Develop programming expertise in Python or Go, focusing on creating automation and tooling for production-grade environments. ‱ Demonstrate expertise in programming with Python and/or Go, coupled with experience creating production-grade automation, tooling, and platform services. ‱ Influence cross-team technical decisions, mentor engineers, elevate technical standards, and collaborate effectively with product engineering teams. ‱ Gain experience in AI/ML infrastructure, model deployment, or GPU workloads to enhance technical expertise and practical understanding. ‱ Design reliability into innovative platforms at the system level while building influence with product engineering teams through technical expertise.

đŸ–ïž Benefits

‱ Your health ‱ Your finances ‱ Your family ‱ Your time at work ‱ Your time pursuing other endeavors

Apply Now

Similar Jobs

🕒 March 24

TechTorch

51 - 200

đŸ’Œ Consulting

📩 Logistics

📣 Marketing

DevOps Engineer at TechTorch managing primary AWS cloud infrastructure and supporting Azure/GCP integrations. Focused on building secure, scalable platforms for enterprise clients.

AWS

Azure

Cloud

Google Cloud Platform

Terraform

🕒 March 24

IT.HR | Recruitment Agency

11 - 50

đŸ’Œ Consulting

📣 Marketing

📩 Logistics

Senior DevOps Engineer specializing in Azure and Kubernetes at leading tech provider. Designing cloud-native architectures, building CI/CD pipelines, and mentoring teams in a remote setup.

Ansible

Azure

Cloud

Grafana

Kubernetes

Prometheus

Python

Terraform

Vault

🕒 March 24

IT.HR | Recruitment Agency

11 - 50

đŸ’Œ Consulting

📣 Marketing

📩 Logistics

Senior DevOps Engineer managing cloud infrastructure for AI and Data workloads at a leading technology provider. Focused on high-availability, scalable cloud-native architectures and automation.

Airflow

Apache

AWS

Azure

Cloud

Google Cloud Platform

Grafana

Kubernetes

Prometheus

Python

Spark

Terraform

🕒 March 20

Alter Solutions Portugal

501 - 1000

đŸ’Œ Consulting

📩 Logistics

đŸ„ Healthcare

DevOps Engineer optimizing CI/CD pipelines and build systems for global IT clients. Collaborate with C++ developers to enhance workflows and infrastructure.

Linux

Python

Terraform

🕒 March 20

Mirantis

501 - 1000

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

DevOps engineer designing and deploying cloud infrastructure products using Kubernetes at Mirantis. Collaborating in a global team to automate and troubleshoot performance issues.

Ansible

AWS

Docker

Firewalls

Jenkins

Kubernetes

Linux

OpenStack

Python

Go