ML Evaluation Engineer

TEKsystems

Singapore

On-site

SGD 90,000 - 150,000

Full time

21 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Allegis Group Singapore Pte Ltd is seeking an ML Evaluation Engineer/Data Scientist to bring statistical rigor and measurement excellence to AI and LLM evaluations. You will design analyses, ensure reliability and reproducibility of results, and partner with ML teams to answer critical questions about performance and trust.

Responsibilities include developing sampling methods, conducting significance tests, building dashboards, and providing data-driven recommendations on evaluation quality and

Qualifications

  • Bachelor's degree in Statistics, Data Science, CS, Mathematics, ML, or related field.
  • Minimum 2+ years in Data Science, Evaluation Science, Applied Statistics, or related disciplines.
  • Strong foundation in statistical analysis, hypothesis testing, experimental design, and measurement methodologies.
  • Experience with structured and unstructured data using Python, SQL, and data analysis tools.
  • Ability to communicate statistical findings clearly to technical and non-technical stakeholders.
  • Experience building analytical dashboards and reporting solutions.
  • Master's or PhD in Statistics, Data Science, CS, or related field a plus.
  • Experience supporting LLM, AI, NLP, or ML evaluation initiatives is a bonus.

Responsibilities

  • Design and execute statistical analyses for AI and LLM evaluation programs.
  • Measure and improve inter-annotator agreement between human evaluators.
  • Develop sampling methodologies and evaluation datasets to ensure representative results.
  • Perform significance testing and statistical validation of model performance changes.
  • Analyze golden sets and benchmark datasets to assess model quality and consistency.
  • Build dashboards and reporting frameworks to monitor evaluation outcomes and trends.
  • Identify sources of evaluation noise, bias, and inconsistency.
  • Partner with ML Evaluation Engineers to validate automated judges against human raters.
  • Provide data-driven recommendations regarding evaluation quality and model readiness.

Skills

Bachelor's degree in Statistics
2+ years of Data Science experience
Statistical analysis
Python & SQL
Communication of statistics
Dashboards & reporting

Education

Bachelor's degree in Statistics, Data Science, CS, Mathematics, ML
Master's or PhD preferred

Tools

Python
SQL
Data analysis tools

Job description

Summary

We are looking for an ML Evaluation Engineer/Data Scientist to bring statistical rigor and measurement excellence to the evaluation of AI and LLM systems. This role will focus on ensuring that evaluation results are reliable, reproducible, and scientifically valid.


The successful candidate will partner closely with ML Evaluation Engineers and Applied Scientists to answer critical questions such as: "Is this performance improvement real?", "Are our evaluators aligned?", and "Can we trust these results?"


Key Responsibilities


  • Design and execute statistical analyses for AI and LLM evaluation programs.

  • Measure and improve inter-annotator agreement between human evaluators.

  • Develop sampling methodologies and evaluation datasets to ensure representative results.

  • Perform significance testing and statistical validation of model performance changes.

  • Analyze golden sets and benchmark datasets to assess model quality and consistency.

  • Build dashboards and reporting frameworks to monitor evaluation outcomes and trends.

  • Identify sources of evaluation noise, bias, and inconsistency.

  • Partner with ML Evaluation Engineers to validate automated judges against human raters.

  • Provide data-driven recommendations regarding evaluation quality and model readiness.


Must Have Skills


  • Bachelor's degree in Statistics, Data Science, Computer Science, Mathematics, Machine Learning, or a related field.

  • Minimum 2+ years of professional experience in Data Science, Evaluation Science, Applied Statistics, or related disciplines.

  • Strong foundation in statistical analysis, hypothesis testing, experimental design, and measurement methodologies.

  • Experience working with structured and unstructured data using Python, SQL, and data analysis tools.

  • Ability to communicate statistical findings clearly to technical and non-technical stakeholders.

  • Experience building analytical dashboards and reporting solutions.


Nice To Haves


  • Master's degree or PhD in Statistics, Data Science, Computer Science, Quantitative Social Sciences, or a related discipline.

  • Direct experience supporting LLM, AI, NLP, or ML evaluation initiatives.

  • Experience measuring agreement between human raters and automated evaluation systems.

  • Knowledge of annotation quality frameworks, golden-set construction, sampling theory, and evaluation benchmarking.

  • Familiarity with LLM-as-Judge frameworks and AI evaluation platforms.


We regret to inform that only shortlisted candidates will be notified / contacted.


EA registration number : YAP JIA YI, R25157934


Allegis Group Singapore Pte Ltd, Company Reg No. 200909448N, EA Licence No. 10C4544

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Evaluation Engineer
ML Evaluation Engineer

ALLEGIS GROUP SINGAPORE PRIVATE LIMITED • Singapore

On-site
SGD 90,000 - 150,000
ML Evaluation Engineer - Statistical QA for AI/LLMs
ML Evaluation Engineer - Statistical QA for AI/LLMs

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 70,000 - 110,000
AI Specialist - TEKsystems (Allegis Group Singapore Pte Ltd)
AI Specialist - TEKsystems (Allegis Group Singapore Pte Ltd)

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
ML Evaluation Engineer: Rigorous AI Benchmarking
ML Evaluation Engineer: Rigorous AI Benchmarking

TEKsystems • Singapore

On-site
SGD 90,000 - 150,000
ML Evaluation Scientist - Data-Driven AI QA
ML Evaluation Scientist - Data-Driven AI QA

ALLEGIS GROUP SINGAPORE PRIVATE LIMITED • Singapore

On-site
SGD 90,000 - 150,000
AI Engineer (Evaluation)
AI Engineer (Evaluation)

Nanyang Technological University Singapore • Singapore

On-site
SGD 90,000 - 130,000
AI Evaluation Scientist: Statistics-Driven Model Benchmarks
AI Evaluation Scientist: Statistics-Driven Model Benchmarks

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 90,000 - 130,000
AI Evaluation Architect
AI Evaluation Architect

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
AI Product Analyst (Developer Community), AI Verify Foundation
AI Product Analyst (Developer Community), AI Verify Foundation

sggovterp • Singapore

On-site
SGD 60,000 - 90,000
AI Product Analyst (Developer Community), AI Verify Foundation
AI Product Analyst (Developer Community), AI Verify Foundation

IMDA • Singapore

On-site
SGD 60,000 - 90,000