LLM Evaluation Engineer – Benchmarking & Automation

UMELIFE (SINGAPORE) PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

UMELIFE (SINGAPORE) PTE. LTD. seeks an AI evaluation engineer to build and maintain an automated LLM evaluation pipeline covering general capabilities, agent capabilities, and persona/role-playing evaluations with one-click assessment and regression testing.

You will execute benchmarks such as MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, manage results across training runs, and prepare checkpoint reports to guide model development.

Qualifications

  • Bachelor's degree or above in Computer Science, AI, or related field.
  • Familiarity with mainstream LLM evaluation benchmarks and frameworks such as lm-eval-harness, OpenCompass, and EvalPlus.
  • Strong proficiency in Python, with the ability to build evaluation pipelines for model inference/deployment, batch evaluation, and results analysis.
  • Familiarity with LLM inference frameworks such as vLLM and SGLang, with the ability to deploy models for batch inference and evaluation.
  • Experience in evaluation data analysis and visualization.
  • Detail-oriented and rigorous, with focus on reproducibility and reliability of evaluation results.

Responsibilities

  • Build and maintain an automated LLM evaluation pipeline covering multiple dimensions.
  • Conduct general capability evaluations using benchmarks like MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval.
  • Conduct Agent capability evaluations and track metrics for BFCL, τ-bench, GAIA.
  • Design persona/role-playing evaluation frameworks, covering identity recognition, role compatibility, multi-turn stability, and style consistency.
  • Record and analyze evaluation results from training runs and produce checkpoint evaluation reports.
  • Conduct regular intermediate evaluations during pre-training to track evolution of capabilities.

Skills

Python
LLM evaluation frameworks familiarity
Data analysis and visualization
Detail-oriented

Education

Bachelor's degree in Computer Science/AI or related field

Tools

lm-eval-harness
OpenCompass
EvalPlus
vLLM
SGLang
CI/CD integration

Job description

UMELIFE (SINGAPORE) PTE. LTD. seeks an AI evaluation engineer to build and maintain an automated LLM evaluation pipeline covering general capabilities, agent capabilities, and persona/role-playing evaluations with one-click assessment and regression testing.

You will execute benchmarks such as MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, manage results across training runs, and prepare checkpoint reports to guide model development.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Evaluation Engineer: Benchmarking & Persona Testing
LLM Evaluation Engineer: Benchmarking & Persona Testing

KUAILU SOFTWARE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 70,000 - 120,000
LLM Evaluation Engineer: Automated Benchmarks & Analytics
LLM Evaluation Engineer: Automated Benchmarks & Analytics

Kuailu Software • Singapore

On-site
SGD 90,000 - 150,000
ML Evaluation Engineer - Statistical QA for AI/LLMs
ML Evaluation Engineer - Statistical QA for AI/LLMs

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 70,000 - 110,000
Large Language Model (LLM) Evaluation Engineer
Large Language Model (LLM) Evaluation Engineer

KUAILU SOFTWARE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 70,000 - 120,000
ML Evaluation Engineer: Rigorous AI Benchmarking
ML Evaluation Engineer: Rigorous AI Benchmarking

TEKsystems • Singapore

On-site
SGD 90,000 - 150,000
Large Language Model (LLM) Evaluation Engineer
Large Language Model (LLM) Evaluation Engineer

Kuailu Software • Singapore

On-site
SGD 90,000 - 150,000
ML Evaluation Engineer
ML Evaluation Engineer

TEKsystems • Singapore

On-site
SGD 90,000 - 150,000
Large Language Model (LLM) Evaluation Engineer
Large Language Model (LLM) Evaluation Engineer

UMELIFE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Evaluation Scientist: Statistics-Driven Model Benchmarks
AI Evaluation Scientist: Statistics-Driven Model Benchmarks

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 90,000 - 130,000
LLM Evaluation Engineer – Multilingual AI Metrics
LLM Evaluation Engineer – Multilingual AI Metrics

Nanyang Technological University • Singapore

On-site
SGD 90,000 - 130,000