Large Language Model (LLM) Evaluation Engineer

KUAILU SOFTWARE (SINGAPORE) PTE. LTD.

Singapore

On-site

SGD 70,000 - 120,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

KUAILU SOFTWARE (SINGAPORE) PTE. LTD. seeks a capable engineer to build and maintain an automated LLM evaluation pipeline across general, agent, and persona settings, enabling one-click evaluation and regression testing.

You will run benchmarks like MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval; set up evaluation environments, track metrics, analyze results, and produce reports to guide pre-training progress.

Qualifications

  • Bachelor's degree or higher in CS/AI or related field.
  • Familiar with LLM evaluation benchmarks and frameworks.
  • Strong Python proficiency for building evaluation pipelines.
  • Experience with model inference frameworks such as vLLM.
  • Experience in data analysis and visualization.
  • Commitment to reproducibility and reliability of results.

Responsibilities

  • Build and maintain an automated LLM evaluation pipeline across general, agent, and persona dimensions.
  • Conduct general capability evaluations using benchmarks such as MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval.
  • Conduct Agent capability evaluations and track benchmark metrics.
  • Design persona/role-playing evaluation frameworks covering identity recognition, role compatibility, multi-turn stability, and style consistency.
  • Record and analyze evaluation results from training runs, conduct comparative analyses and anomaly detection, and produce checkpoint evaluation reports.
  • Conduct regular intermediate evaluations during the pre-training stage to track evolution and improvement of model capabilities.

Skills

Python
LLM evaluation
Benchmarking
CI/CD automation
Data analysis

Education

Bachelor's degree in Computer Science or AI

Tools

lm-eval-harness
OpenCompass
EvalPlus
vLLM
SGLang

Job description

Job Responsibilities
  1. Build and maintain an automated LLM evaluation pipeline covering multiple dimensions, including general capabilities, Agent capabilities, and persona/role-playing. The pipeline should support one-click evaluation, historical result comparison, and regression testing.
  2. Conduct general capability evaluations using benchmarks such as MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, including benchmark deployment, execution, and results analysis.
  3. Conduct Agent capability evaluations, including setting up evaluation environments and tracking metrics for benchmarks such as BFCL, τ-bench, and GAIA.
  4. Design and execute persona/role-playing evaluation frameworks, covering metrics such as identity recognition, role compatibility, multi-turn stability, and style consistency.
  5. Record and analyze evaluation results from training runs, conduct comparative analysis and anomaly detection, and produce checkpoint evaluation reports.
  6. Conduct regular intermediate evaluations during the pre‑training stage to track the evolution and improvement of model capabilities.
Job Requirements
  1. Bachelor's degree or above in Computer Science, Artificial Intelligence, or a related field.
  2. Familiarity with mainstream LLM evaluation benchmarks and frameworks, such as lm‑eval‑harness, OpenCompass, and EvalPlus.
  3. Strong proficiency in Python, with the ability to independently build evaluation pipelines covering model inference/deployment, batch evaluation, and results analysis.
  4. Familiarity with LLM inference frameworks such as vLLM and SGLang, with the ability to deploy models for batch inference and evaluation.
  5. Experience in evaluation data analysis and visualization.
  6. Detail‑oriented and rigorous, with a strong focus on ensuring the reproducibility and reliability of evaluation results.
Preferred Qualifications
  1. Experience with Agent evaluation, particularly BFCL, τ-bench, GAIA, or SWE‑bench.
  2. Experience with persona or role‑playing evaluation, such as CharacterBench or RMTBench.
  3. Experience with evaluation automation and CI/CD integration.
  4. Understanding of model training workflows, with the ability to understand the relationship between training checkpoints and evaluation results.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Large Language Model (LLM) Evaluation Engineer
Large Language Model (LLM) Evaluation Engineer

Kuailu Software • Singapore

On-site
SGD 90,000 - 150,000
Large Language Model (LLM) Evaluation Engineer
Large Language Model (LLM) Evaluation Engineer

UMELIFE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
LLM Evaluation Engineer: Automated Benchmarks & Analytics
LLM Evaluation Engineer: Automated Benchmarks & Analytics

Kuailu Software • Singapore

On-site
SGD 90,000 - 150,000
LLM Evaluation Engineer – Benchmarking & Automation
LLM Evaluation Engineer – Benchmarking & Automation

UMELIFE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
LLM Evaluation Engineer: Benchmarking & Persona Testing
LLM Evaluation Engineer: Benchmarking & Persona Testing

KUAILU SOFTWARE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 70,000 - 120,000
ML Evaluation Engineer
ML Evaluation Engineer

TEKsystems • Singapore

On-site
SGD 90,000 - 150,000
AI Engineer – LLM Algorithm Engineer (Agentic Commerce)
AI Engineer – LLM Algorithm Engineer (Agentic Commerce)

SHOPEE IP SINGAPORE PRIVATE LIMITED • Singapore

On-site
SGD 180,000 - 240,000
Sr. AI/ML Engineer
Sr. AI/ML Engineer

GECO Asia Pte Ltd • Singapore

On-site
SGD 150,000 - 210,000
AI Engineer
AI Engineer

BASIL TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 180,000 - 280,000
Senior AI Engineer (Principal-Level Scope)
Senior AI Engineer (Principal-Level Scope)

CHEMT BIOTECHNOLOGY PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000