Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Kuailu Software is seeking an experienced evaluation engineer to build and maintain automated LLM evaluation pipelines. You will cover general, agent, and persona-based benchmarks, enabling one-click evaluation, historical comparisons, and regression testing.
You will deploy and run benchmarks like MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, while tracking metrics for BFCL, τ-bench, and GAIA. Strong Python and evaluation framework experience are essential.
Kuailu Software is seeking an experienced evaluation engineer to build and maintain automated LLM evaluation pipelines. You will cover general, agent, and persona-based benchmarks, enabling one-click evaluation, historical comparisons, and regression testing.
You will deploy and run benchmarks like MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, while tracking metrics for BFCL, τ-bench, and GAIA. Strong Python and evaluation framework experience are essential.