Senior Site Reliability Engineer

Job not on LinkedIn

🕒 April 22

🌐 Egypt, Jordan – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đŸ‘» Ghost score 24%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Unifonic

Unifonic

501 - 1000 employees

Founded 2006

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

💰 $125M Series B on 2021-09

Consulting ‱ Healthcare ‱ Logistics

Unifonic is a customer engagement platform powered by artificial intelligence, designed to facilitate personalized omnichannel communication. The company offers a variety of communication channels and applications, including SMS, Voice, WhatsApp, Push Notifications, and Webchat to enhance customer interactions across multiple industries. Unifonic focuses on improving marketing automation, IT and operations, and customer support by providing AI-powered tools and automated workflows, ensuring timely and efficient customer communication. Their solutions cater to industries such as retail, banking, healthcare, and logistics, providing integrations with popular platforms like Salesforce and Shopify. With over 25 billion messages sent and 5,000+ customers, Unifonic provides global connectivity and a seamless customer experience.

📋 Description

‱ Owning the reliability, uptime, and scalability of critical production services 24/7. ‱ Participating in the on-call rotation to respond to incidents, troubleshoot live production issues, and lead post-incident analysis. ‱ Building robust operational playbooks, escalation paths, and improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). ‱ Ensuring operational excellence by proactively detecting and addressing reliability risks through SLO monitoring, chaos testing, and capacity planning. ‱ Automating operational tasks to minimize human intervention. ‱ Architecting, implementing, and managing infrastructure across AWS, Oracle Cloud Infrastructure (OCI), and OpenStack environments. ‱ Optimizing cloud resources to balance performance, security, and cost-efficiency. ‱ Managing Kubernetes clusters (EKS, OKE, Rancher RKE2) for scalability, availability, and performance. ‱ Managing and optimizing high-performance messaging and caching systems including Kafka, RabbitMQ, and Redis. ‱ Managing and optimizing production-grade MySQL and PostgreSQL databases. ‱ Leading the planning and execution of comprehensive disaster recovery strategies. ‱ Implementing advanced observability solutions (Prometheus, Grafana, CloudWatch). ‱ Driving automation initiatives using Terraform, Helm, Jenkins, Tekton or GitLab CI/CD. ‱ Integrating security best practices into infrastructure and applications. ‱ Collaborating with cross-functional teams to foster SRE culture and mentoring junior engineers.

🎯 Requirements

‱ Bachelor's or master's degree in computer science, Engineering, or a related technical field. ‱ 8+ years of hands-on production experience in SRE, DevOps, or cloud engineering roles. ‱ Strong expertise in AWS, OCI, OpenStack environments. ‱ Deep understanding of Kubernetes ecosystems (EKS, OKE, Rancher RKE2). ‱ Proven experience with Kafka, RabbitMQ, Redis, and distributed messaging and caching systems. ‱ Solid experience managing MySQL and PostgreSQL in production environments. ‱ Expert-level scripting and automation skills (Python, Bash, Go). ‱ Advanced proficiency with Helm, Terraform, and modern CI/CD toolchains. ‱ Demonstrable experience with Linux system administration and troubleshooting. ‱ Must be available at night during the on-call schedule.

đŸ–ïž Benefits

‱ Competitive salary and bonus ‱ Unifonic share scheme (we are all owners!) ‱ 30 holiday days after the first anniversary ‱ Your Birthday off! ‱ Spend up to 25 days per year working from anywhere in the world! ‱ Paid leave and assistance for new parents ‱ LinkedIn learning license

Apply Now