Senior AI Backend Engineer – Agent Evaluation, Quality

🔥 12 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Salla E-Commerce Platform

Salla E-Commerce Platform

51 - 200 employees

مستقبل التجارة الإلكترونية، ابدأ تجارتك بسهولة 🛒

📋 Description

• Own the evaluation stack. Design and build LLM-as-judge systems, calibrate them against human labels, and make agent quality measurable per-agent and per-failure-mode. • Make the release gate real. Build per-PR eval harnesses and regression detection wired into CI, so quality is enforced automatically, not by manual passes. • Build user simulators to generate test coverage and adversarial cases before real users hit them. • Turn production signal into improvement - pipe real failures back into evaluation sets so the system compounds over time. • Partner with product to turn "what good looks like" into concrete, measurable criteria. • Grow into agent development - contribute to building and hardening the agents themselves, starting with the components you know most deeply from evaluating them.

🎯 Requirements

• Strong software engineering fundamentals. Production Python or Typescript (or similar), clean API and system design, testing, CI/CD. You write code others build on - evaluation infrastructure is real engineering. • Hands-on LLM/agent experience. You've built with LLMs - agents, RAG, tool/function calling, orchestration frameworks (LangGraph, LangChain, or equivalent) - and understand how they behave and break. • A measurement mindset. You reason about metrics, calibration, and experiments; you want to quantify whether something works, not just ship it. • Production experience. You've run LLM systems in production and dealt with reliability, latency, cost, and observability. • 5+ years software engineering, with recent hands-on LLM/agent work. • Nice to have • Direct experience evaluating LLM/agent systems - offline/online eval, LLM-as-judge, systematic regression testing. • Observability tooling (Arize, LangSmith, or similar). • Arabic language / NLP experience. • E-commerce or merchant-facing product experience.

Apply Now

Similar Jobs

🕒 April 1

InnovationTeam

201 - 500

🏢 Enterprise

☁️ SaaS

🔌 API

Quality Assurance Engineer ensuring software quality and functionality at InnovationTeam. Collaborating with teams on web and mobile applications for a forward-thinking technology company.

🇸🇦 Saudi Arabia – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🔧 QA Engineer (Quality Assurance)

Selenium

SQL

TFS

🕒 February 5

Tabby

201 - 500

💳 Fintech

🛍️ eCommerce

Senior QA Engineer designing testing strategy and performing hands-on testing for a global fintech company. Collaborating with an international engineering team for product reliability.

🇸🇦 Saudi Arabia – Remote

💰 $58M Series C on 2023-01

⏰ Full Time

🟠 Senior

🔧 QA Engineer (Quality Assurance)

Android

Cloud

Docker

Grafana

iOS

Kubernetes

Microservices

Prometheus

TypeScript

Go