Senior AI Backend Engineer - Agent Evaluation & Quality

Salla

Makkah Region

On-site

SAR 300,000 - 540,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Salla is seeking a senior engineer to own the evaluation stack for its production multi-agent systems. You will design LLM-as-judge components, calibrate against human labels, and quantify agent quality per failure mode.

You will also build simulators, integrate regression detection into CI, and contribute to agent development by turning failures into improvements. This role emphasizes strong software engineering and production readiness.

Qualifications

  • 5+ years software engineering with recent hands-on LLM/agent work.
  • Experience building evaluation stacks, judges, and simulators.
  • Strong fundamentals in Python or TypeScript, APIs, testing and CI/CD.
  • Hands-on experience with LangGraph, LangChain, or equivalent.

Responsibilities

  • Own the evaluation stack and design LLM-as-judge systems.
  • Integrate per-PR eval harnesses and regression detection into CI.
  • Build user simulators to generate test coverage and adversarial cases.
  • Turn real failures into improved evaluation data and criteria.
  • Collaborate with product to formalize measurable success criteria.

Skills

LLM/agent experience
Measurement mindset
Production experience
System design

Tools

Python
TypeScript
LangGraph
LangChain
CI/CD
APIs

Job description

About the role

We run production multi-agent systems that handle real work for a large base of users. As those systems grow, our biggest constraint is confidence: we need to know how well the agents perform, catch regressions before they ship, and keep quality steady as we release. This role owns that.

You’ll build the evaluation systems behind our agents – the judges, test harnesses, and simulators that tell us whether an agent is working and where it’s failing. The goal is to let us ship agents faster because we can trust what the evaluation tells us.

Evaluation is the focus, but it won’t be the boundary. Because you’ll understand the agents’ failure modes better than anyone there will also be opportunities to contribute to agent development itself, building and improving the agents alongside the systems that evaluate them.

Responsibilities
  • Own the evaluation stack. Design and build LLM-as-judge systems, calibrate them against human labels, and make agent quality measurable per‑agent and per‑failure‑mode.
  • Make the release gate real. Build per‑PR eval harnesses and regression detection wired into CI, so quality is enforced automatically, not by manual passes.
  • Build user simulators to generate test coverage and adversarial cases before real users hit them.
  • Turn production signal into improvement – pipe real failures back into evaluation sets so the system compounds over time.
  • Partner with product to turn “what good looks like” into concrete, measurable criteria.
  • Grow into agent development – contribute to building and hardening the agents themselves, starting with the components you know most deeply from evaluating them.
  • Strong software engineering fundamentals. Production Python or Typescript (or similar), clean API and system design, testing, CI/CD. You write code others build on – evaluation infrastructure is real engineering.
  • Hands‑on LLM/agent experience. You’ve built with LLMs – agents, RAG, tool/function calling, orchestration frameworks (LangGraph, LangChain, or equivalent) – and understand how they behave and break.
  • A measurement mindset. You reason about metrics, calibration, and experiments; you want to quantify whether something works, not just ship it.
  • Production experience. You’ve run LLM systems in production and dealt with reliability, latency, cost, and observability.
  • 5+ years software engineering, with recent hands‑on LLM/agent work.
Nice to have
  • Direct experience evaluating LLM/agent systems – offline/online eval, LLM-as-judge, systematic regression testing.
  • Observability tooling (Arize, LangSmith, or similar).
  • Arabic language / NLP experience.
  • E‑commerce or merchant‑facing product experience.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Backend Engineer — Agent Evaluation & Quality
Senior AI Backend Engineer — Agent Evaluation & Quality

Salla • Makkah Region

On-site
SAR 300,000 - 540,000
AI Engineer
AI Engineer

Saudi Azm عزم السعودية • Riyadh

On-site
SAR 260,000 - 460,000
AI Engineer
AI Engineer

Latitude • Riyadh

On-site
SAR 420,000 - 660,000
Expert / Director of Al
Expert / Director of Al

P2P • Makkah Region

On-site
SAR 450,000 - 750,000
AI Systems Architect - LLM & Vector Infrastructure
AI Systems Architect - LLM & Vector Infrastructure

Starmarkets • Riyadh

On-site
SAR 300,000 - 520,000
Forward Deployed Engineer, Agentic Platform
Forward Deployed Engineer, Agentic Platform

Tkxel • Riyadh

On-site
SAR 300,000 - 500,000
Forward Deployed Engineer, Agentic Platform
Forward Deployed Engineer, Agentic Platform

Tkxel LLC • Riyadh

On-site
SAR 250,000 - 350,000
Software Engineer with Applied AI Experience
Software Engineer with Applied AI Experience

Tkxel • Riyadh

On-site
SAR 250,000 - 420,000
Agentic AI Engineer
Agentic AI Engineer

Webook • Saudi Arabia

On-site
SAR 180,000 - 320,000
Agentic AI, Engineer
Agentic AI, Engineer

Master Works • Riyadh

On-site
SAR 240,000 - 360,000