Agent Evaluation Engineer for an AI Agent Company

DADACONSULTANTS PTE. LTD.

Singapore

On-site

SGD 90,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Frontier AI team
Unlimited AI tools access
Competitive salary

Job summary

DADACONSULTANTS PTE. LTD. is seeking a role focused on evaluating AI agents and models post-training, developing metrics, tasks and experiments to measure performance and value.

You will design evaluation criteria, analyze results, and collaborate with product, engineering and model teams to validate improvements, ensuring robust and unbiased insights for next-gen AI systems.

Qualifications

  • Evidence of metric design and validation.
  • Understanding of AI model and agent mechanisms.
  • Experience in data analysis, experiment design and handling uncertainty.
  • Ability to explain findings and their limitations clearly.

Responsibilities

  • Design online and offline metrics translating user tasks and output quality into measurable criteria.
  • Create evaluation tasks and experiments around agent planning, tool use, context and feedback.
  • Support post-training evaluations and compare third-party model quality.
  • Develop new evaluation methods for missed issues and test validity and bias.
  • Analyze evaluation results and work with product, engineering and model teams to validate improvements.

Skills

Metric design
Experiment design
Data analysis
Agent mechanisms
Task planning

Job description

About our client

Our client is a fast-growing, well-funded AI agent company. The product takes on complex work like deep research, data analysis and software development, and carries it through to a finished result. It is a small team building at the frontier, based in Singapore.

About the role

You will develop evaluations grounded in product needs, user tasks and how AI models and agents actually work. Using metrics, experiments and failure analysis, you'll assess whether capability changes are real, investigate gaps in existing evaluations, and inform system development, post-training and model selection.

What you'll do
  • Design online and offline metrics that translate user tasks, output quality and practical value into measurable, testable evaluation criteria
  • Design evaluation tasks and experiments around agent planning, tool use, context and feedback, comparing performance before and after system changes
  • Support post-training evaluations and comparisons of third-party model quality, defining use cases and the conditions results apply to
  • Develop new tasks, metrics or experimental methods for issues existing evaluations miss, and test their validity, bias and reproducibility
  • Analyze evaluation results and failures, distinguish score changes from real capability changes, and work with product, engineering and model teams to validate improvements
What you bring
  • Practical experience evaluating agent products or model post-training, with concrete evidence of metric design and validation
  • Deep understanding of AI model and agent mechanisms — task planning, tool use, context management and feedback
  • Strong data analysis, engineering and research skills, including experiment design and handling uncertainty in results
  • Ability to investigate open-ended problems, develop well-reasoned new evaluation methods, and examine experimental bias
  • Independent judgement on AI output quality and real user value, with the ability to explain findings and their limitations clearly
Nice to have
  • Familiarity with user research methods (interviews, observation, usability testing), and the ability to translate findings into evaluation criteria
What we offer
  • Build at the frontier of AI agents with a fast-paced team
  • Unlimited access to our own AI tools
  • Competitive salary, reviewed for the right person
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Agent Evaluation Engineer
Agent Evaluation Engineer

Manus AI • Singapore

On-site
SGD 90,000 - 150,000
Equity plan
Unlimited Manus Tokens
Frontier AI projects
AI Agent Evaluation Architect
AI Agent Evaluation Architect

Manus AI • Singapore

On-site
SGD 90,000 - 150,000
Equity plan
Unlimited Manus Tokens
Frontier AI projects
Senior AI Evaluation Engineer #AIDA
Senior AI Evaluation Engineer #AIDA

Singtel • Singapore

On-site
Confidential
Principal AI Engineer
Principal AI Engineer

Resaro • Singapore

On-site
SGD 180,000 - 240,000
Agent Harness Engineer (Agent Runtime & Infrastructure) | AI Agent Company
Agent Harness Engineer (Agent Runtime & Infrastructure) | AI Agent Company

DADACONSULTANTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 200,000
Research Engineer, Benchmarks
Research Engineer, Benchmarks

Clera • Singapore

On-site
SGD 190,000 - 317,000
Visa sponsorship available
Senior AI Agent Researcher — Evaluation-Driven Systems
Senior AI Agent Researcher — Evaluation-Driven Systems

XG TECH PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Agent Engineer
AI Agent Engineer

GoHire Technologies LTD. • Singapore

On-site
SGD 120,000 - 180,000
Applied ML / AI Engineer
Applied ML / AI Engineer

Nestria Ai Northstar Ai Shield Pte. Ltd. • Singapore

On-site
SGD 70,000 - 100,000
Senior Technical Product Manager - AI Agents, Evals & Reliability
Senior Technical Product Manager - AI Agents, Evals & Reliability

AI Chopping Block • Singapore

On-site
SGD 120,000 - 180,000