Principal Site Reliability Engineer

đŸ”„ 0 minutes ago

đŸ‡”đŸ‡± Poland – Remote

⏰ Full Time

🔮 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

đŸ‘» Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Akamai Technologies

Akamai Technologies

5001 - 10000 employees

🔒 Cybersecurity

💰 Post-IPO Equity on 2001-07

Cloud Computing ‱ Cybersecurity ‱ Content Delivery

Akamai Technologies is a leading cloud services provider that specializes in delivering security, cloud computing, and content delivery solutions. It offers a range of services such as API security, DDoS protection, and performance optimization for web applications, ensuring secure and reliable user experiences. With a robust global infrastructure, Akamai empowers businesses to streamline their digital presence while safeguarding against various cyber threats and enhancing application performance.

📋 Description

‱ Shape compute-platform strategy and qualification criteria for x64, ARM, accelerator, and inference hardware across firmware, OS, runtime, and reliability ‱ Lead provisioning and CI/CD improvements, including reproducible, safe, and scalable bare-metal provisioning, imaging, configuration, and infrastructure delivery pipelines ‱ Drive reliability and observability by addressing systemic failures, enhancing telemetry and diagnostics, and establishing safe production rollout and recovery safeguards ‱ Investigate complex hardware/software incidents, identify root causes, guide decision-making, and convert findings into engineering fixes ‱ Review designs, mentor engineers, establish standards, and resolve multi-domain issues across teams and vendors ‱ Enable new hardware through provisioning services and provide seamless firmware upgrade models ‱ Ensure reliable server behavior across Akamai datacenters

🎯 Requirements

‱ Expertise in Linux systems and server platforms, including boot processes, storage, networking, hardware diagnostics, firmware, BIOS/UEFI, BMCs, and server lifecycles ‱ Experience designing and operating bare-metal provisioning, configuration management, and scalable infrastructure automation ‱ Fluency in Python and Bash ‱ Experience building automation, APIs, deployment pipelines, and debugging multi-layer failures ‱ Ability to identify whether hardware-test failures originate from firmware, BIOS/UEFI configuration, device firmware, kernel/driver behavior, or the physical test environment ‱ Expertise in observability, metrics, capacity analysis, incident root-cause investigation, and production readiness ‱ Practical knowledge of x86 and ARM platforms, accelerators, and inference infrastructure, including drivers, runtimes, and compatibility ‱ Hands-on experience with Linux virtualization technologies, including KVM, QEMU, and libvirt ‱ Understanding of nested virtualization, CPU virtualization extensions, VM networking and storage, and Linux-level troubleshooting of virtualized environments

đŸ–ïž Benefits

‱ Health, well-being, finances, and life beyond work benefits ‱ FlexBase workplace flexibility: work at home, in an office, or a combination of both

Apply Now

Similar Jobs

🕒 September 10

Commit

501 - 1000

🔒 Cybersecurity

Principal DevOps Engineer building AWS, Kubernetes, and delivery infrastructure for a regulated gaming platform. Scaling reliable sportsbook operations for millions of players and peak live-betting traffic.

AWS

Cloud

EC2

Grafana

Kafka

Kubernetes

Prometheus

Python

Terraform

Vault

Go

🕒 June 7

Madiff

51 - 200

đŸ’Œ Consulting

📩 Logistics

đŸ„ Healthcare

Staff DevOps Engineer responsible for driving cloud platform architecture and infrastructure strategy in a remote environment. Leading architectural discussions while mentoring engineers across multiple teams.

AWS

Cloud

Grafana

Jenkins

Kubernetes

Linux

Prometheus

SDLC

Terraform