LLM Evaluation Engineer: Benchmarking & Persona Testing

KUAILU SOFTWARE (SINGAPORE) PTE. LTD.

Singapore

On-site

SGD 70,000 - 120,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

KUAILU SOFTWARE (SINGAPORE) PTE. LTD. seeks a capable engineer to build and maintain an automated LLM evaluation pipeline across general, agent, and persona settings, enabling one-click evaluation and regression testing.

You will run benchmarks like MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval; set up evaluation environments, track metrics, analyze results, and produce reports to guide pre-training progress.

Qualifications

  • Bachelor's degree or higher in CS/AI or related field.
  • Familiar with LLM evaluation benchmarks and frameworks.
  • Strong Python proficiency for building evaluation pipelines.
  • Experience with model inference frameworks such as vLLM.
  • Experience in data analysis and visualization.
  • Commitment to reproducibility and reliability of results.

Responsibilities

  • Build and maintain an automated LLM evaluation pipeline across general, agent, and persona dimensions.
  • Conduct general capability evaluations using benchmarks such as MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval.
  • Conduct Agent capability evaluations and track benchmark metrics.
  • Design persona/role-playing evaluation frameworks covering identity recognition, role compatibility, multi-turn stability, and style consistency.
  • Record and analyze evaluation results from training runs, conduct comparative analyses and anomaly detection, and produce checkpoint evaluation reports.
  • Conduct regular intermediate evaluations during the pre-training stage to track evolution and improvement of model capabilities.

Skills

Python
LLM evaluation
Benchmarking
CI/CD automation
Data analysis

Education

Bachelor's degree in Computer Science or AI

Tools

lm-eval-harness
OpenCompass
EvalPlus
vLLM
SGLang

Job description

KUAILU SOFTWARE (SINGAPORE) PTE. LTD. seeks a capable engineer to build and maintain an automated LLM evaluation pipeline across general, agent, and persona settings, enabling one-click evaluation and regression testing.

You will run benchmarks like MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval; set up evaluation environments, track metrics, analyze results, and produce reports to guide pre-training progress.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Evaluation Engineer: Automated Benchmarks & Analytics
LLM Evaluation Engineer: Automated Benchmarks & Analytics

Kuailu Software • Singapore

On-site
SGD 90,000 - 150,000
LLM Evaluation Engineer – Benchmarking & Automation
LLM Evaluation Engineer – Benchmarking & Automation

UMELIFE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Large Language Model (LLM) Evaluation Engineer
Large Language Model (LLM) Evaluation Engineer

KUAILU SOFTWARE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 70,000 - 120,000
Large Language Model (LLM) Evaluation Engineer
Large Language Model (LLM) Evaluation Engineer

Kuailu Software • Singapore

On-site
SGD 90,000 - 150,000
Large Language Model (LLM) Evaluation Engineer
Large Language Model (LLM) Evaluation Engineer

UMELIFE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
ML Evaluation Engineer - Statistical QA for AI/LLMs
ML Evaluation Engineer - Statistical QA for AI/LLMs

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 70,000 - 110,000
ML Evaluation Engineer
ML Evaluation Engineer

TEKsystems • Singapore

On-site
SGD 90,000 - 150,000
ML Evaluation Engineer: Rigorous AI Benchmarking
ML Evaluation Engineer: Rigorous AI Benchmarking

TEKsystems • Singapore

On-site
SGD 90,000 - 150,000
AI Evaluation Scientist: Statistics-Driven Model Benchmarks
AI Evaluation Scientist: Statistics-Driven Model Benchmarks

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 90,000 - 130,000
LLM Evaluation Engineer – Multilingual AI Metrics
LLM Evaluation Engineer – Multilingual AI Metrics

Nanyang Technological University • Singapore

On-site
SGD 90,000 - 130,000