Una candidatura hecha a medida para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.
Kuailu Software is seeking an experienced evaluation engineer to build and maintain automated LLM evaluation pipelines. You will cover general, agent, and persona-based benchmarks, enabling one-click evaluation, historical comparisons, and regression testing.
You will deploy and run benchmarks like MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, while tracking metrics for BFCL, τ-bench, and GAIA. Strong Python and evaluation framework experience are essential.