Senior Evaluation Algorithm Engineer

Binance

Hong Kong

Hybrid

HKD 600,000 - 1,000,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Work-from-home arrangement
Competitive salary and benefits
Career development opportunities

Job summary

Binance is seeking an evaluation expert with an algorithmic background to build LLM evaluation capabilities spanning dialogue, trading, and other business scenarios. You will quantify model performance, pinpoint issues, and drive continuous model improvement within a fast-paced, global fintech environment.

The role emphasizes designing robust evaluation plans, constructing high-quality datasets, and scaling evaluation workflows to support rapid model iterations while collaborating with multiple

Qualifications

  • Master's degree or above in CS/AI/Math/Stats with solid algorithmic foundation.
  • Hands-on LLM evaluation experience at a large tech company involved in commercial deployment.
  • Familiar with evaluation methods and able to define dimensions for different business scenarios.
  • Systematic control over evaluation data representativeness, annotation consistency, and result reliability.
  • Proficient in Python and automation of evaluation workflows.
  • Strong business understanding and cross-team communication.

Responsibilities

  • Design end-to-end LLM evaluation plans for dialogue and financial trading scenarios.
  • Lead dataset construction with clear rubrics and quality control processes.
  • Analyze results to identify failures and provide actionable improvement recommendations.
  • Develop scalable evaluation platforms and toolchains for rapid model iteration.
  • Collaborate with algorithm, product, and data teams to translate objectives into evaluation standards and drive RD directions.

Skills

Python
LLM evaluation
Data analysis
Cross-functional collaboration
Communication

Education

Master's degree in CS/AI/Math/Stats

Tools

Evaluation platforms
Benchmark construction
Python automation

Job description

Binance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. We are trusted by 300+ million people in 100+ countries for our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-asset products. Binance offerings range from trading and finance to education, research, payments, institutional services, Web3 features, and more. We leverage the power of digital assets and blockchain to build an inclusive financial ecosystem to advance the freedom of money and improve financial access for people around the world.

About the Role

In the AI era, large language models are reshaping core business scenarios such as dialogue and trading. Model capability iteration relies on a scientific and trustworthy evaluation system — the "ruler" that measures model quality and guides R&D direction. We are seeking an evaluation expert with an algorithmic background to build LLM evaluation capabilities covering dialogue, financial trading, and other scenarios, using professional evaluation methods to quantify model performance, pinpoint issues, and drive continuous model improvement.

Responsibilities
  1. Design end-to-end LLM evaluation plans for business scenarios such as dialogue and financial trading. Build evaluation metric systems and rubrics, transforming subjective model performance judgments into quantifiable, reproducible, and explainable evaluation conclusions.
  2. Lead the design and construction of evaluation datasets. Define evaluation dimensions and scenario coverage, establish high-quality data annotation guidelines and quality control processes, and build benchmarks that authentically reflect business needs and have discriminative power.
  3. Analyze model capability boundaries and failure modes based on evaluation results. Produce actionable improvement recommendations and collaborate with algorithm and product teams to drive model iteration, making evaluation a critical component of the R&D loop.
  4. Drive the automation and scaling of evaluation workflows. Build sustainable evaluation platforms and toolchains to support high-frequency, stable evaluation needs during rapid model iteration.
  5. Collaborate with algorithm, product, and data teams to translate business and model objectives into clear evaluation standards, and turn evaluation findings into concrete R&D directions and drive their implementation.
Requirements
  1. Master's degree or above in Computer Science, Artificial Intelligence, Mathematics, Statistics, or related fields, with a solid algorithmic foundation and understanding of LLM principles, training, and fine-tuning processes.
  2. Hands-on LLM evaluation experience at a large tech company, with participation in commercial deployment evaluation (not purely academic or offline benchmarking). Familiar with the full pipeline from evaluation data preparation and rubrics design to evaluation-driven R&D.
  3. Familiar with mainstream evaluation methods (human evaluation, model-based automatic evaluation / LLM-as-a-judge, metric computation) and their applicable boundaries. Able to define appropriate evaluation dimensions for different business scenarios and write clear, actionable, and discriminative rubrics.
  4. Systematic control over evaluation data representativeness, annotation consistency, and result reliability, ensuring scientific and trustworthy evaluation conclusions.
  5. Proficient in Python, with experience in evaluation workflow automation, benchmark construction, or evaluation platform development. Able to independently handle data processing, evaluation script writing, and result analysis.
  6. Strong business understanding and communication skills, able to translate evaluation findings into clear improvement directions and effectively drive cross-team collaboration.
Bonus Qualifications
  1. Experience evaluating dialogue systems, AI Agents, or financial/trading LLMs.
  2. Experience building high-quality AI training/evaluation data or data annotation systems.
  3. Familiarity with RLHF, reward models, or preference data-related work.
Why Binance
  • Shape the future with the world’s leading blockchain ecosystem
  • Collaborate with world-class talent in a user-centric global organization with a flat structure
  • Tackle unique, fast-paced projects with autonomy in an innovative environment
  • Thrive in a results-driven workplace with opportunities for career growth and continuous learning
  • Competitive salary and company benefits
  • Work-from-home arrangement (the arrangement may vary depending on the work nature of the business team)

Binance is committed to being an equal opportunity employer. We believe that having a diverse workforce is fundamental to our success.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Evaluation Engineer — Metrics & Automation
Senior LLM Evaluation Engineer — Metrics & Automation

Binance • Hong Kong

On-site
HKD 600,000 - 1,000,000
Work-from-home arrangement
Competitive salary and benefits
Career development opportunities
Senior Large Language Model Algorithm Engineer/Expert
Senior Large Language Model Algorithm Engineer/Expert

Binance • Hong Kong

On-site
HKD 861,393 - 1,096,319
Binance Accelerator Program - Research Data Scientist
Binance Accelerator Program - Research Data Scientist

Binance • Hong Kong

On-site
HKD 313,479 - 548,589
Competitive salary
Company benefits
Work-from-home arrangement
Algorithm Engineer, Market Growth
Algorithm Engineer, Market Growth

Binance • Hong Kong Island

On-site
HKD 600,000 - 800,000
Competitive salary
Company benefits
Work-from-home arrangement
+2
Pioneer Talent Program - Applied Data Scientist
Pioneer Talent Program - Applied Data Scientist

Binance • Hong Kong

On-site
HKD 548,159 - 783,085
Competitive salary
Work-from-home arrangements
Opportunities for career growth
Binance Accelerator Program - Backend Engineer (AI Pro / Agent Infrastructure) Fully Remote
Binance Accelerator Program - Backend Engineer (AI Pro / Agent Infrastructure) Fully Remote

Binance • Hong Kong

On-site
HKD 20,088 - 27,900
Stipend
Career exposure
Business Intelligence/ Data Analytics Manager
Business Intelligence/ Data Analytics Manager

Lever, Inc. • Hong Kong

Remote
HKD 300,000 - 600,000
Binance Accelerator Program - Applied AI Agent Developer
Binance Accelerator Program - Applied AI Agent Developer

Binance • Hong Kong

On-site
HKD 167,400 - 279,000
Competitive salary
Work-from-home arrangement
Career growth opportunities
AI Engineer (LLM/ Chatbot)
AI Engineer (LLM/ Chatbot)

Pantheon Lab Limited • Hong Kong

On-site
HKD 900,000 - 1,300,000
Binance Accelerator Program - AI Productivity Engineer
Binance Accelerator Program - AI Productivity Engineer

Binance • Hong Kong

On-site
HKD 44,640 - 66,960
Competitive salary
Company benefits
Work-from-home arrangement