LLM Evaluation Engineer: Automated Benchmarks & Analytics

Kuailu Software

Singapore

On-site

SGD 90,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Kuailu Software is seeking an experienced evaluation engineer to build and maintain automated LLM evaluation pipelines. You will cover general, agent, and persona-based benchmarks, enabling one-click evaluation, historical comparisons, and regression testing.

You will deploy and run benchmarks like MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, while tracking metrics for BFCL, τ-bench, and GAIA. Strong Python and evaluation framework experience are essential.

Qualifications

  • Bachelor's degree or above in Computer Science, Artificial Intelligence, or a related field.
  • Familiarity with mainstream LLM evaluation benchmarks and frameworks, such as lm-eval-harness, OpenCompass, and EvalPlus.
  • Strong proficiency in Python, with the ability to independently build evaluation pipelines covering model inference/deployment, batch evaluation, and results analysis.

Responsibilities

  • Build and maintain an automated LLM evaluation pipeline covering multiple dimensions, including general capabilities, Agent capabilities, and persona/role-playing. The pipeline should support one-click evaluation, historical result comparison, and regression testing.
  • Conduct general capability evaluations using benchmarks such as MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, including benchmark deployment, execution, and results analysis.
  • Conduct Agent capability evaluations, including setting up evaluation environments and tracking metrics for benchmarks such as BFCL, τ-bench, and GAIA.
  • Design and execute persona/role-playing evaluation frameworks, covering metrics such as identity recognition, role compatibility, multi-turn stability, and style consistency.
  • Record and analyze evaluation results from training runs, conduct comparative analysis and anomaly detection, and produce checkpoint evaluation reports.
  • Conduct regular intermediate evaluations during the pre-training stage to track the evolution and improvement of model capabilities.

Skills

Python
LLM evaluation
Benchmark frameworks
CI/CD integration
Data visualization

Education

Bachelor's degree in Computer Science / AI

Tools

lm-eval-harness
OpenCompass
EvalPlus
vLLM
SGLang

Job description

Kuailu Software is seeking an experienced evaluation engineer to build and maintain automated LLM evaluation pipelines. You will cover general, agent, and persona-based benchmarks, enabling one-click evaluation, historical comparisons, and regression testing.

You will deploy and run benchmarks like MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, while tracking metrics for BFCL, τ-bench, and GAIA. Strong Python and evaluation framework experience are essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Evaluation Engineer: Benchmarking & Persona Testing
LLM Evaluation Engineer: Benchmarking & Persona Testing

KUAILU SOFTWARE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 70,000 - 120,000
LLM Evaluation Engineer – Benchmarking & Automation
LLM Evaluation Engineer – Benchmarking & Automation

UMELIFE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Large Language Model (LLM) Evaluation Engineer
Large Language Model (LLM) Evaluation Engineer

KUAILU SOFTWARE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 70,000 - 120,000
Large Language Model (LLM) Evaluation Engineer
Large Language Model (LLM) Evaluation Engineer

Kuailu Software • Singapore

On-site
SGD 90,000 - 150,000
Large Language Model (LLM) Evaluation Engineer
Large Language Model (LLM) Evaluation Engineer

UMELIFE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
ML Evaluation Engineer: Rigorous AI Benchmarking
ML Evaluation Engineer: Rigorous AI Benchmarking

TEKsystems • Singapore

On-site
SGD 90,000 - 150,000
ML Evaluation Engineer - Statistical QA for AI/LLMs
ML Evaluation Engineer - Statistical QA for AI/LLMs

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 70,000 - 110,000
ML Evaluation Engineer
ML Evaluation Engineer

TEKsystems • Singapore

On-site
SGD 90,000 - 150,000
Senior LLM Performance & Evaluation Engineer
Senior LLM Performance & Evaluation Engineer

Bitdeer Group • Singapore

On-site
SGD 120,000 - 180,000
AI Evaluation Scientist: Statistics-Driven Model Benchmarks
AI Evaluation Scientist: Statistics-Driven Model Benchmarks

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 90,000 - 130,000