Senior Site Reliability Engineer – Cloud and Networking

🕒 May 28

🇵🇱 Poland – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 23%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Akamai Technologies

Akamai Technologies

5001 - 10000 employees

🔒 Cybersecurity

💰 Post-IPO Equity on 2001-07

Cloud Computing • Cybersecurity • Content Delivery

Akamai Technologies is a leading cloud services provider that specializes in delivering security, cloud computing, and content delivery solutions. It offers a range of services such as API security, DDoS protection, and performance optimization for web applications, ensuring secure and reliable user experiences. With a robust global infrastructure, Akamai empowers businesses to streamline their digital presence while safeguarding against various cyber threats and enhancing application performance.

📋 Description

• Owning the SRE lifecycle for NodeBalancer and Network Load Balancer — from design reviews and pre-rollout readiness assessments through production sign-off and ongoing reliability management • Designing and implementing SLO/SLI frameworks that reflect true customer experience for L4 and L7 load balancing services, and driving action when error budgets are at risk • Building and maintaining observability pipelines for NB/NLB infrastructure, including Prometheus metrics from load balancing components and system-level sources, and Grafana dashboards that enable rapid incident triage • Leading technical incident response for complex NB/NLB failures — BGP/VIP issues, failover failures, data plane degradations, and configuration problems — acting as the technical commander and driving root cause analysis and preventive follow-through • Developing and automating safe deployment workflows for phased NB/NLB releases, including bake period monitoring, feature flag management, and GO/NO-GO validation across global datacenter rollouts • Reviewing design documents, product requirement Documents and producing actionable SRE input on operational risks, capacity implications, Day-2 concerns, and product strategy gaps • Building automation and tooling using Python or Go that reduces operational toil and improves team-wide operational capability • Mentoring SRE II engineers on the NB team, providing hands-on technical guidance, code/config reviews, and raising the bar for the team's SRE practice • Participating in an on-call rotation for NB/NLB production systems, responding to incidents and driving resolution for customer-facing load balancing infrastructure • Participate in a scheduled, daytime-only on-call rotation to spearhead technical incident response and resolve complex NB/NLB failures.

🎯 Requirements

• Have extensive experience in SRE, platform engineering, or infrastructure engineering, working with large-scale distributed systems • Demonstrate deep expertise with Linux networking fundamentals — routing, BGP, nftables/iptables, ARP, VXLAN — and comfort diagnosing at the packet level using tcpdump, netstat, and similar tools • Have hands-on experience with L4/L7 load balancing technologies — including proxy-based or kernel-level load balancers — covering configuration, health checking, high availability, and failure modes at scale • Show a track record of defining SLO/SLI frameworks, building observability platforms from scratch, and running incident management processes at scale • Demonstrate expertise in Kubernetes and containerization at scale — including workload scheduling, networking (CNI, Services, ingress), resource management, and operating stateful or network-intensive workloads in a cluster environment • Build automation and tooling using Python or Go, with infrastructure-as-code experience (SaltStack, Ansible, or Terraform) and strong deployment safety instincts • Demonstrate 4+ years in SRE or infrastructure engineering, with at least 2 years at cloud scale

🏖️ Benefits

• Your health • Your finances • Your family • Your time at work • Your time pursuing other endeavors

Apply Now

Similar Jobs

🕒 May 27

Inetum

10,000+ employees

💼 Consulting

🏥 Healthcare

🛡️ Insurance

System Engineer enhancing VoIP and Radius systems at Inetum Polska. Operating in high-performance data centers and driving project lifecycles from design to deployment.

Linux

MariaDB

MySQL

OpenStack

Perl

Puppet

Python

TCP/IP

VoIP

🕒 May 23

Madiff

51 - 200

💼 Consulting

📦 Logistics

🏥 Healthcare

DevOps Engineer supporting a cloud-based EGM management platform across complex infrastructures. Focusing on automation, CI/CD, and improving developer experience in a gaming technology environment.

AWS

Cloud

Jenkins

Kubernetes

Python

SDLC

Terraform

🕒 May 20

Sigma Software Group

1001 - 5000

💼 Consulting

🏥 Healthcare

🚘 Automotive

AI Deployment Engineer designing and optimizing AI-driven solutions. Collaborate with clients to build and deploy effective AI systems for operational environments.

🕒 May 19

Netguru

501 - 1000

💼 Consulting

🏥 Healthcare

📣 Marketing

Forward Deployment Engineer at Netguru delivering AI systems for clients. Engaging in client delivery and internal AI tooling with a focus on rapid prototyping.

🗣️🇵🇱 Polish Required

Cloud

Python

SDLC

TypeScript

🕒 May 19

Inetum

10,000+ employees

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Senior DevOps Engineer focusing on producing a scalable machine learning solution for churn prevention. Collaborating with engineering teams to implement Kubernetes and CI/CD best practices.

Airflow

Ansible

Apache

Docker

Grafana

Kubernetes

Prometheus

Terraform

Vault