Senior AI Evaluation Engineer

Singtel

Singapore

On-site

Confidential

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Singtel is seeking a Senior Technical Evaluator to lead the design, calibration, and adjudication of evaluation suites and gate thresholds across agent archetypes. You will own regression-pack methodologies for AI/ML model changes and serve as the technical custodian of evaluation quality.

You will design offline evaluation suites, conduct operability gate reviews, calibrate production thresholds, and produce monthly quality reports.

Qualifications

  • 6+ years in ML/data/software with strong evaluation or quality focus.
  • Hands-on experience evaluating LLM or ML systems.
  • LLM/agent evaluation design and statistical rigour.
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Clear technical writing for gate findings.
  • Ability to influence build teams on quality.

Responsibilities

  • Design and maintain offline evaluation suites (golden sets, regression packs, adversarial probes) and the continuous-evaluation scoring pipeline across archetypes.
  • Conduct operability gate reviews: assess evidence packs, reproduce results, and recommend go/no-go with documented findings.
  • Own the model-update regression pack methodology and adjudicate regression runs against archived baselines.
  • Calibrate gate thresholds against production reality and maintain evaluation drift hygiene.
  • Produce the monthly quality report per agent: eval trends, defect clusters with reproduction traces.
  • Mentor team members and review their work for quality and consistency.

Skills

ML/DS experience
LLM/agent evaluation
Python & evaluation tooling
Data analysis & metrics
Tracing & observability
Analytical rigor
Technical writing
Quality influence
Telco customer intents

Education

Bachelor’s or Master’s in Computer Science or related field

Tools

promptfoo
DeepEval
custom evaluation harness

Job description

Serve as the senior technical evaluator within the team, responsible for the design, calibration, and adjudication of evaluation suites and gate thresholds across agent archetypes. Lead gate reviews under the Team Lead’s authority, own the regression-pack methodology for AI/ML model changes, and act as the technical custodian of evaluation quality and drift hygiene.

Make an Impact by:

  • Design and maintain offline evaluation suites (golden sets, regression packs, adversarial/safety probes) and the continuous-evaluation scoring pipeline across archetypes.
  • Conduct operability gate reviews: assess evidence packs, reproduce evaluation results, and recommend go/no-go with documented findings.
  • Own the model-update regression pack methodology and adjudicate regression runs against archived baselines with AIML Operations team.
  • Calibrate gate thresholds against production reality and maintain evaluation drift hygiene (golden-set rotation, hold-out sets, judge calibration).
  • Produce the monthly quality report per agent: eval trends, failure-mode taxonomy, and defect clusters with reproduction traces.
  • Mentor members of the team when needed and review their work for quality and consistency.

Skills for Success:

  • Bachelor’s or Master’s degree in Computer Science or a related field
  • 6+ years in ML/data/software with strong evaluation or quality focus
  • Hands-on experience evaluating LLM or ML systems
  • LLM/agent evaluation design and statistical rigour
  • Python and evaluation tooling (promptfoo, DeepEval, or custom harnesses)
  • Data analysis and metric interpretation
  • Tracing/observability tooling
  • Analytical rigour and attention to detail
  • Clear technical writing for gate findings
  • Ability to influence build teams on quality
  • Understanding of telco customer intents and journeys
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Lead & Gate Architect
AI Evaluation Lead & Gate Architect

Singtel • Singapore

On-site
Confidential
AI Evaluation Engineer
AI Evaluation Engineer

ALLEGIS GROUP SINGAPORE PRIVATE LIMITED • Singapore

On-site
SGD 120,000 - 160,000
AI Evaluation Engineer: Optimize & Validate LLMs
AI Evaluation Engineer: Optimize & Validate LLMs

ALLEGIS GROUP SINGAPORE PRIVATE LIMITED • Singapore

On-site
SGD 120,000 - 160,000
Lead AI Engineer
Lead AI Engineer

Airwallex • Singapore

On-site
SGD 180,000 - 300,000
#EG AI Engineer
#EG AI Engineer

NCS Group • Singapore

On-site
SGD 80,000 - 120,000
AI Specialist - TEKsystems (Allegis Group Singapore Pte Ltd)
AI Specialist - TEKsystems (Allegis Group Singapore Pte Ltd)

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
AI Systems Engineer (Model Routing)
AI Systems Engineer (Model Routing)

SignalPlus • Singapore

On-site
SGD 180,000 - 240,000
AI Engineer
AI Engineer

Accenture Southeast Asia • Singapore

On-site
SGD 180,000 - 240,000
Senior AI Harness Engineer
Senior AI Harness Engineer

AVENSYS CONSULTING PTE. LTD. • Singapore

On-site
SGD 150,000 - 190,000
Senior Specialist, AI Engineer
Senior Specialist, AI Engineer

AIA Hong Kong and Macau • Singapore

On-site
SGD 80,000 - 120,000