Research Engineer, AI Evaluation & Agent Reliability

Mollkom

Riyadh

Hybrid

SAR 180,000 - 360,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Mollkom in Riyadh is seeking a specialist in ML evaluation to design benchmarks for catalog, content, conversations and tool-use tasks, and to develop regression and adversarial tests.

You will analyze experiments, communicate findings clearly, and translate insights into engineering improvements. This hybrid role is based in Riyadh, Saudi Arabia.

Qualifications

  • Experience in ML evaluation, applied research or AI quality engineering.
  • Strong Python, data analysis and experiment design.
  • Understanding of LLM-as-judge limitations, data leakage and evaluation bias.
  • Ability to communicate results clearly and translate findings into engineering improvements.

Responsibilities

  • Build benchmarks for catalog, content, conversations and tool-use tasks.
  • Design rubrics and human/automated evaluation with agreement, bias and coverage analysis.
  • Develop regression and adversarial tests for permissions and unsupported claims.
  • Analyze experiments and model comparisons, communicating actionable findings and statistical limits.

Skills

Python
Data analysis
Experiment design
LLM evaluation

Job description

Commerce outcomes require more than fluent answers. Design evaluations that distinguish good writing, correct information and completed action, helping the team release improvements with measurable effects.

Responsibilities
  • Build benchmarks for catalog, content, conversations and tool-use tasks.
  • Design rubrics and human/automated evaluation with agreement, bias and coverage analysis.
  • Develop regression and adversarial tests for permissions and unsupported claims.
  • Analyze experiments and model comparisons, communicating actionable findings and statistical limits.
Qualifications
  • Experience in ML evaluation, applied research or AI quality engineering.
  • Strong Python, data analysis and experiment design.
  • Understanding of LLM-as-judge limitations, data leakage and evaluation bias.
  • Ability to communicate results clearly and translate findings into engineering improvements.

Hybrid in Riyadh, Saudi Arabia.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Backend Engineer - Agent Evaluation & Quality
Senior AI Backend Engineer - Agent Evaluation & Quality

Salla • Makkah Region

On-site
SAR 300,000 - 540,000
Senior AI Backend Engineer - Agent Evaluation & Quality
Senior AI Backend Engineer - Agent Evaluation & Quality

Salla • Saudi Arabia

On-site
SAR 300,000 - 520,000
Staff AI Engineer, Agent Systems
Staff AI Engineer, Agent Systems

Mollkom • Riyadh

Hybrid
SAR 240,000 - 420,000
Senior AI Engineer, Retrieval & Knowledge Systems
Senior AI Engineer, Retrieval & Knowledge Systems

Mollkom • Riyadh

Hybrid
SAR 240,000 - 360,000
Senior Product Engineer, AI Experiences
Senior Product Engineer, AI Experiences

Mollkom • Riyadh

Hybrid
SAR 167,000 - 391,000
Agentic AI Engineer
Agentic AI Engineer

Webook • Saudi Arabia

On-site
SAR 180,000 - 320,000
Agent Engineer
Agent Engineer

Sarj | سرج • Riyadh

On-site
SAR 150,000 - 210,000
null
Senior AI Backend Engineer: Agent Evaluation & Quality
Senior AI Backend Engineer: Agent Evaluation & Quality

Salla • Saudi Arabia

On-site
SAR 300,000 - 520,000
AI Engineer
AI Engineer

Saudi Azm عزم السعودية • Riyadh

On-site
SAR 260,000 - 460,000
AI Agent Engineer: Build & QA Real-World Agents
AI Agent Engineer: Build & QA Real-World Agents

Sarj • Riyadh

On-site
SAR 120,000 - 180,000