Senior Site Reliability Engineer

🕒 May 11

đŸ‡”đŸ‡± Poland – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đŸ‘» Ghost score 52%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Akamai Technologies

Akamai Technologies

5001 - 10000 employees

🔒 Cybersecurity

🏱 Enterprise

đŸ“± Media

Cybersecurity ‱ Enterprise ‱ Media

Akamai Technologies is a global edge platform and cloud services company that delivers content delivery, edge computing, and security solutions. The company operates one of the world’s largest distributed networks to accelerate and protect web, media, and application traffic, offering products for content delivery, DDoS protection, API and app security, bot management, edge compute (serverless/edge functions), and AI inference at the edge. Akamai also provides enterprise-focused security services (zero trust, identity and access management, secure internet access) and cloud/AI infrastructure tools, and has recently expanded capabilities through acquisitions (for example LayerX) to add browser-based AI usage control.

📋 Description

‱ Providing support and mentorship for other engineers within the department ‱ Developing and maintaining automated tools and scripts to enhance system reliability, deployment processes, and incident response efficiency. ‱ Improving our system monitoring to speed error detection and remediation, enhancing performance and reliability of virtualization platform ‱ Participating in on-call rotations, guiding restoration and repair of service-impacting issues ‱ Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response ‱ Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure

🎯 Requirements

‱ Possess expert level experience in a SysAdmin (Linux/Unix Administration), DevOps or SRE role, working with large scale distributed systems ‱ Demonstrate expertise in Kubernetes and large-scale containerization systems. ‱ Possess at least one programming language (Python/Golang) and configuration management with Terraform/SaltStack/Ansible ‱ Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring. ‱ Have experience with architecting software and infrastructure at scale ‱ Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices.

đŸ–ïž Benefits

‱ Your health ‱ Your finances ‱ Your family ‱ Your time at work ‱ Your time pursuing other endeavors

Apply Now

Similar Jobs

🕒 April 23

RedSky

11 - 50

đŸ’Œ Consulting

đŸŽ–ïž Defense

đŸ„ Healthcare

Venture Builder creating startups from the ground up at Red Sky. Join and build teams pushing boundaries across various industries.

🕒 April 22

CloudLinux

51 - 200

☁ SaaS

🔐 Security

🌐 Web 3

Lead the evolution of CloudLinux's data platform into a DBaaS model. Design resilient databases and implement automated infrastructure management for high-performance systems.

Airflow

Ansible

Apache

Cloud

ETL

Kubernetes

MongoDB

Postgres

Python

Redis

SQL

Terraform

Zookeeper

Go

🕒 April 20

VirtusLab

201 - 500

đŸ’Œ Consulting

🏭 Manufacturing

📩 Logistics

DevOps Engineer managing security operations for a rapidly scaling UK insurance leader. Responsible for incident response, security analysis, and integrating cloud infrastructure for insurance solutions.

Cloud

🕒 April 15

Madiff

51 - 200

đŸ’Œ Consulting

📩 Logistics

đŸ„ Healthcare

Site Reliability Engineer responsible for monitoring and managing AI workloads in a remote environment. Ensuring system stability and improving operational excellence for AI-driven systems.

Azure

Grafana

Kubernetes

🕒 April 2

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Senior SRE responsible for automation, architecture decisions, and managing AI workloads for Akamai. Collaborating with product teams on reliability and operational readiness.

Distributed Systems

Grafana

Kubernetes

Prometheus

Python

Terraform

Go