Site Reliability Engineer

🕒 March 31

🇸🇦 Saudi Arabia – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 40%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Lucidya | لوسيديا

Lucidya | لوسيديا

51 - 200 employees

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

💰 $6M Series B on 2022-01

Consulting • Marketing • Artificial Intelligence

Lucidya is a company specializing in social media and online platform analytics. They provide tools and insights that help businesses understand customer sentiment, track brand reputation, and make data-driven decisions. Lucidya's solutions are tailored for engaging with the Arabic-speaking market, leveraging AI technology to process and analyze data for optimal marketing strategies.

📋 Description

• Design and maintain highly available, fault-tolerant, and scalable infrastructure • Identify and eliminate single points of failure • Ensure production systems remain stable under increasing scale and load • Manage and continuously improve workloads across AWS, GCP, or Azure • Use Terraform to standardize and scale infrastructure • Optimize resource usage for performance and cost • Operate and scale production Kubernetes clusters, including EKS and GKE • Troubleshoot Kubernetes issues and ensure smooth deployments and upgrades • Ensure containerized workloads perform reliably at scale • Implement and refine monitoring with Prometheus, Grafana, Datadog, or ELK • Define meaningful alerting • Respond to incidents, lead root cause analysis, and apply lessons learned • Write scripts and build tooling to eliminate repetitive operational work • Improve infrastructure efficiency through automation • Collaborate with DevOps and engineering teams on performance bottlenecks • Contribute to CI/CD improvements and deployment reliability • Shape reliability best practices across the organization • Build understanding of infrastructure, systems, and workflows during the first 30 days • Independently manage infrastructure tasks and troubleshooting by 90 days • Take ownership of infrastructure areas and improve reliability and scalability

🎯 Requirements

• ~3 years of experience in SRE, DevOps, or infrastructure engineering • Experience working in cloud environments such as AWS, GCP, or Azure • Understanding of distributed systems • Hands-on production experience with Kubernetes • Experience troubleshooting Kubernetes issues • Experience using Terraform or similar Infrastructure as Code tools • Proficiency with Docker and Kubernetes • Scripting experience in Python, Bash, or similar • Understanding of CI/CD pipelines, including Jenkins, GitHub Actions, or Bitbucket • Understanding of networking, load balancing, and high-availability design • Experience implementing monitoring tools such as Prometheus, Grafana, Datadog, or ELK • Ability to distinguish useful alerts from noise • Ownership, calmness under pressure, methodical incident response, clear communication, and ability to simplify complexity • Experience with RabbitMQ or Redis in production is nice to have, but not required • Familiarity with Ansible or AWX is nice to have, but not required • Exposure to multi-cloud or hybrid environments is nice to have, but not required • AWS, GCP, or Linux certifications are nice to have, but not required • Background from ITI is nice to have, but not required

Apply Now