Site Reliability Engineer – Guardicore AI Platform

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Akamai Technologies

Akamai Technologies

5001 - 10000 employees

🔒 Cybersecurity

💰 Post-IPO Equity on 2001-07

Cloud Computing • Cybersecurity • Content Delivery

Akamai Technologies is a leading cloud services provider that specializes in delivering security, cloud computing, and content delivery solutions. It offers a range of services such as API security, DDoS protection, and performance optimization for web applications, ensuring secure and reliable user experiences. With a robust global infrastructure, Akamai empowers businesses to streamline their digital presence while safeguarding against various cyber threats and enhancing application performance.

📋 Description

• Operating secure, highly available Kubernetes infrastructure for core microservices, data pipelines, observability, and internal tooling. • Enhancing platform reliability, observability, security, performance, and cost efficiency. • Providing guidance to engineers and developers to increase confidence that their services are performing as expected. • Leading complex production investigations and driving long-term improvements. • Leverage LLMs and AI-driven automation to auto-remediate incidents and streamline operations. • Partner across DevOps, Software, Data, AI and Security engineering Teams to investigate and troubleshoot complex problems. • Participating in on-call rotations, guiding restoration and repair of service-impacting issues.

🎯 Requirements

• 3+ years of experience in SRE, DevOps, or Platform Engineering, with a proven track record of mastering and troubleshooting complex system architectures. • Demonstrate ability to design and implement a comprehensive monitoring and observability strategy using tools like Prometheus and Grafana. • Have production experience with Kubernetes, Docker, Helm, and third-party clouds (GCP, Azure, Linode, AWS) on Linux-based infrastructure. • Have exceptional troubleshooting and problem-solving skills across network, system, applications, and database layers. • Have experience with GitOps, CI/CD, and Infrastructure as Code. • Have scripting and programming proficiency in Python, Go, and Bash. • Leverage AI tools in daily operational tasks and actively propose initiatives to improve platform automation. • Demonstrate technical leadership and ownership in driving cross-team initiatives, defining tools, and building foundational frameworks.

🏖️ Benefits

• We support your health, well-being, finances, and life beyond work. • FlexBase adapts to your job's needs • It's about supporting employees to do their best work. We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Apply Now

Similar Jobs

🔥 4 hours ago

As a DevOps / Platform Engineer, you'll drive automation and cloud operations for a secure platform. Join a global company providing converged security solutions for over 30 years.

Ansible

Azure

Cloud

Distributed Systems

Docker

Grafana

Kafka

Linux

Prometheus

🕒 2 days ago

Fundraise Up

51 - 200

🤲 Charity

💳 Fintech

☁️ SaaS

Senior DevOps Engineer for Fundraise Up, handling CI/CD, observability, and mentor teams. Join a global nonprofit platform enhancing donation experiences for nonprofits worldwide.

🗣️🇷🇺 Russian Required

Ansible

Docker

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Python

🕒 2 days ago

IRIUM

501 - 1000

🔒 Cybersecurity

☁️ SaaS

Cloud DevOps Engineer specializing in Microsoft environments for a full-remote project with IRIUM. Seeking candidates with Azure experience and strong English skills.

🗣️🇪🇸 Spanish Required

Azure

Docker

Kubernetes

🕒 5 days ago

IRIUM

501 - 1000

🔒 Cybersecurity

☁️ SaaS

DevOps Engineer collaborating on an international project in a full-remote role at IRIUM. Seeking candidates with strong infrastructure knowledge and experience in automation and scripting.

🗣️🇪🇸 Spanish Required

Docker

Groovy

Jenkins

Kubernetes

Linux

Python

🕒 6 days ago

Codurance

51 - 200

☁️ SaaS

🏢 Enterprise

🤝 B2B

Senior Platform Engineer / SRE at Codurance, focusing on cloud migrations and CI/CD setup. Collaborate with agile teams leveraging software craftsmanship and extreme programming practices.

🗣️🇪🇸 Spanish Required

Ansible

AWS

Azure

Chef

Cloud

Grafana

Kubernetes

Prometheus

Puppet

Python

Terraform