Senior AI Backend Engineer – Agent Evaluation, Quality

🕒 July 19

🇸🇦 Saudi Arabia – Remote

⏰ Full Time

🟠 Senior

🔙 Backend Engineer

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Salla E-Commerce Platform

Salla E-Commerce Platform

51 - 200 employees

🛍️ eCommerce

☁️ SaaS

eCommerce • SaaS

Salla E-Commerce Platform is an innovative platform designed to simplify the process of starting and managing online businesses. It specializes in e-commerce solutions, enabling users to create and operate their online shops effortlessly. With a focus on providing accessible tools and services for e-commerce, Salla is shaping the future of online commerce in a user-friendly manner.

📋 Description

• Own the evaluation stack. Design and build LLM-as-judge systems, calibrate them against human labels, and make agent quality measurable per-agent and per-failure-mode. • Make the release gate real. Build per-PR eval harnesses and regression detection wired into CI, so quality is enforced automatically, not by manual passes. • Build user simulators to generate test coverage and adversarial cases before real users hit them. • Turn production signal into improvement - pipe real failures back into evaluation sets so the system compounds over time. • Partner with product to turn "what good looks like" into concrete, measurable criteria. • Grow into agent development - contribute to building and hardening the agents themselves, starting with the components you know most deeply from evaluating them.

🎯 Requirements

• Strong software engineering fundamentals. Production Python or Typescript (or similar), clean API and system design, testing, CI/CD. You write code others build on - evaluation infrastructure is real engineering. • Hands-on LLM/agent experience. You've built with LLMs - agents, RAG, tool/function calling, orchestration frameworks (LangGraph, LangChain, or equivalent) - and understand how they behave and break. • A measurement mindset. You reason about metrics, calibration, and experiments; you want to quantify whether something works, not just ship it. • Production experience. You've run LLM systems in production and dealt with reliability, latency, cost, and observability. • 5+ years software engineering, with recent hands-on LLM/agent work. • Nice to have • Direct experience evaluating LLM/agent systems - offline/online eval, LLM-as-judge, systematic regression testing. • Observability tooling (Arize, LangSmith, or similar). • Arabic language / NLP experience. • E-commerce or merchant-facing product experience.

Apply Now

Similar Jobs

🕒 July 1

Filigran

201 - 500

🔒 Cybersecurity

☁️ SaaS

Customer Platform Architect in cybersecurity, assisting clients with deploying on-premise infrastructure and providing technical support. Collaborating with teams to ensure effective threat management through reliable systems.

🇸🇦 Saudi Arabia – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🔙 Backend Engineer

🗣️🇸🇦 Arabic Required

AWS

Azure

Cloud

Cyber Security

Docker

Google Cloud Platform

Kubernetes

🕒 February 4

TAWANTECH

11 - 50

💳 Fintech

🤝 B2B

Senior Java Backend Developer designing and maintaining scalable Java backend applications. Collaborating remotely with teams to enhance core applications and services.

🇸🇦 Saudi Arabia – Remote

⏰ Full Time

🟠 Senior

🔙 Backend Engineer

AWS

Azure

Cloud

Docker

Hibernate

Java

Kubernetes

Maven

Microservices

MySQL

Oracle

Postgres

Spring

Spring Boot

SpringBoot

SQL