Senior AI Backend Engineer: Agent Evaluation & Quality

Salla

Saudi Arabia

On-site

SAR 300,000 - 520,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Salla is seeking an engineer to own the evaluation stack for its multi-agent platform. You will design and build the judges, test harnesses, and simulators that tell us whether an agent is working and where it is failing.

The role focuses on evaluation yet offers opportunities to contribute to agent development themselves. You will work with product to translate what good looks like into measurable criteria and push quality through CI.

Qualifications

  • Strong software engineering fundamentals.
  • Production Python or TypeScript, clean API and CI/CD.
  • Hands-on LLM/agent experience.
  • Experience with LLMs – agents, RAG, tool calling, orchestration frameworks.
  • A measurement mindset: quantify metrics, calibration, and experiments.
  • Production experience with LLM systems.

Responsibilities

  • Own the evaluation stack and automate it.
  • Design and build LLM-as-judge systems and calibrate against human labels.
  • Make the release gate real.
  • Build per-PR eval harnesses and regression detection wired into CI.
  • Build user simulators to generate test coverage and adversarial cases.
  • Turn production signals into evaluation improvements and feedback loops.
  • Collaborate with product to define concrete, measurable criteria.
  • Grow into agent development alongside evaluation systems.

Skills

Software engineering
Python/Typescript
LLM/Agent work
LLM/Agent systems
Production systems

Tools

LangGraph
LangChain
Arize
LangSmith

Job description

Salla is seeking an engineer to own the evaluation stack for its multi-agent platform. You will design and build the judges, test harnesses, and simulators that tell us whether an agent is working and where it is failing.

The role focuses on evaluation yet offers opportunities to contribute to agent development themselves. You will work with product to translate what good looks like into measurable criteria and push quality through CI.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Backend Engineer — Agent Evaluation & Quality
Senior AI Backend Engineer — Agent Evaluation & Quality

Salla • Makkah Region

On-site
SAR 300,000 - 540,000
Senior AI Backend Engineer - Agent Evaluation & Quality
Senior AI Backend Engineer - Agent Evaluation & Quality

Salla • Makkah Region

On-site
SAR 300,000 - 540,000
Senior AI Backend Engineer - Agent Evaluation & Quality
Senior AI Backend Engineer - Agent Evaluation & Quality

Salla • Saudi Arabia

On-site
SAR 300,000 - 520,000
AI Agent Engineer: Build & QA Real-World Agents
AI Agent Engineer: Build & QA Real-World Agents

Sarj • Riyadh

On-site
SAR 120,000 - 180,000
Research Engineer, AI Evaluation & Agent Reliability
Research Engineer, AI Evaluation & Agent Reliability

Mollkom • Riyadh

Hybrid
SAR 180,000 - 360,000
AI Agent Engineer | Build & QA Voice Agents
AI Agent Engineer | Build & QA Voice Agents

Sarj | سرج • Riyadh

On-site
SAR 150,000 - 210,000
null
Agent Engineer
Agent Engineer

Sarj | سرج • Riyadh

On-site
SAR 150,000 - 210,000
null
Agent Engineer
Agent Engineer

Sarj • Riyadh

On-site
SAR 120,000 - 180,000
AI Agent Engineer Intern - Hands-On, High-Impact
AI Agent Engineer Intern - Hands-On, High-Impact

Sarj • Riyadh

On-site
Hands-on experience in impactful AI products
Opportunity to work in a high-growth, innovative tech start-up
AI Engineering Director — Autonomous Agent Leader
AI Engineering Director — Autonomous Agent Leader

P2P • Makkah Region

On-site
SAR 900,000 - 1,500,000