Senior Site Reliability Engineer

Job not on LinkedIn

šŸ”„ 0 minutes ago

Apply Now
Find Similar Remote Jobs

šŸ“Š Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mozn

Mozn

201 - 500 employees

Founded 2017

šŸ’¼ Consulting

šŸ„ Healthcare

šŸ“¦ Logistics

šŸ’° $10M Series A on 2022-02

Consulting • Healthcare • Logistics

Mozn is a regional AI company building Arabic-native generative AI and enterprise AI platforms. It provides OSOS (an Arabic-first GenAI platform), FOCAL (a financial-crime and fraud detection platform), and customized AI solutions spanning language intelligence, risk intelligence, operational AI, data management, geospatial intelligence, and AI centers. Mozn focuses on serving enterprise customers in the MENA region (including healthcare, finance, and government) with SaaS products and tailored AI services that prioritize cultural relevance, data security, and regulatory compliance.

šŸ“‹ Description

• Carry a normal on-call rotation and act as a hands-on incident responder • Investigate, fix, and document production incidents • Perform application-level debugging and identify root causes in service code and business logic • Ship fixes or pull requests directly into application repositories when appropriate • Design, build, and ship LLM-based agents integrated with Kubernetes, cloud APIs, observability, and incident-management tools • Define agent tool interfaces and build safe wrappers for APIs, scripts, and read/write actions • Establish autonomy guardrails and human-approval requirements for agents • Own agent evaluations and build test/backtest suites using historical incidents • Tune prompts, context, and tool schemas as agent scope expands • Partner with the SRE/platform team to identify suitable automation workflows • Report agent impact using MTTD, MTTR, MTTX, false-positive/negative rates, and engineer-hours of toil removed • Maintain security- and compliance-first operations with audit trails, least-privilege production access, and Saudi data-residency/regulatory alignment

šŸŽÆ Requirements

• 3+ years building production software with LLMs, including agentic workflows, tool/function calling, multi-step planning, or RAG • Hands-on experience shipping work with Claude Code, OpenAI Codex, or Kimi K2/K3 • Strong Python or similar programming skills for agent tooling, API wrappers, and orchestration • Hands-on SRE experience as a primary on-call responder, including incident response and root cause analysis • Application-level debugging skills and ability to read service code, trace failures to underlying logic, and ship fixes • Hands-on Kubernetes and cloud provider experience with AWS, GCP, OCI, or Azure • Fluency with Prometheus, Grafana, Datadog, or ELK • Understanding of autonomous-system guardrails, permissioning, approval gates, rollback paths, and auditability • Ability to build trust with technical stakeholders • Experience in Saudi Arabia/MENA, Terraform/Ansible, Docker, VM/on-prem setups, LLM agent evaluation, ML engineering, LLMOps, or platform engineering are nice to have

šŸ–ļø Benefits

• Competitive compensation • Top-tier health insurance • Responsibility and trust • Freedom and autonomy in decision-making • Fun and dynamic workplace • Opportunity to work alongside leading AI talent • Inclusive and empowering culture • Opportunity to work at the forefront of AI in the Middle East

Apply Now

Similar Jobs

šŸ•’ August 2

Robusta Studio

51 - 200

šŸ’¼ Consulting

šŸ“£ Marketing

šŸ“¦ Logistics

Cloud DevSecOps Engineer responsible for designing, implementing, and securing cloud-based infrastructure and CI/CD pipelines across OCI and GCP environments.

šŸ‡ŖšŸ‡¬ Egypt – Remote

ā° Full Time

🟔 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

Cloud

Google Cloud Platform

Grafana

Kubernetes

Oracle

Prometheus

SDLC

Splunk

Terraform

šŸ•’ July 22

Storyteller

11 - 50

šŸ“£ Marketing

šŸ’¼ Consulting

šŸ“¦ Logistics

Site Reliability Engineer responding to live incidents for a high-growth B2B SaaS platform. Coordinating technical responses and improving reliability for customer systems.

šŸ‡ŖšŸ‡¬ Egypt – Remote

šŸ’µ €20k / year

ā° Full Time

🟔 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Azure

Cloud

šŸ•’ April 23

Unifonic

501 - 1000

šŸ’¼ Consulting

šŸ„ Healthcare

šŸ“¦ Logistics

Senior DevOps Engineer working to enhance CI/CD and integrate systems at Unifonic. Join a dynamic team to shape communication solutions for 5000+ customer-centric companies.

šŸ‡ŖšŸ‡¬ Egypt – Remote

šŸ’° $125M Series B on 2021-09

ā° Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

Cloud

Docker

Kubernetes

Python

Terraform

šŸ•’ April 22

Unifonic

501 - 1000

šŸ’¼ Consulting

šŸ„ Healthcare

šŸ“¦ Logistics

Senior Site Reliability Engineer at Unifonic responsible for enhancing system reliability and managing cloud infrastructure for communication solutions.

šŸ‡ŖšŸ‡¬ Egypt – Remote

šŸ’° $125M Series B on 2021-09

ā° Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

AWS

Cloud

Grafana

Jenkins

Kafka

Kubernetes

Linux

MySQL

OpenStack

Oracle

Postgres

Prometheus

Python

RabbitMQ

Redis

Terraform

Go

šŸ•’ October 23, 2025

Blink

51 - 200

šŸ“£ Marketing

šŸØ Hospitality

āœˆļø Travel

Senior DevOps & Architecture Engineer optimizing CI/CD practices and collaborating with teams remotely. Leading system automation and ensuring scalability and security in operations.

Ansible

AWS

Azure

Chef

Cloud

DNS

Docker

DynamoDB

ElasticSearch

Grafana

HAProxy

Jenkins

Kubernetes

Logstash

MongoDB

MySQL

NGINX

Postgres

Prometheus

Puppet

Python

Redis

SMTP

Terraform

Vault