AI Evaluation Engineer

ALLEGIS GROUP SINGAPORE PRIVATE LIMITED

Singapore

On-site

SGD 120,000 - 160,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Allegis Group Singapore Private Limited is seeking an AI Evaluation Engineer to understand, analyze, and improve the performance of advanced AI systems. You will design evaluation methodologies, analyze model outputs, and translate findings into actionable guidance for technical and business stakeholders.

Responsibilities include leading rigorous evaluations of LLMs and multimodal models, developing scoring frameworks, and building robust MLOps workflows to support scalable, automated

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, ML, AI, or related field.
  • Industry experience in ML engineering or applied research.
  • Advanced proficiency in Python and modern DL ecosystems (PyTorch, JAX, HF).
  • Experience building scalable ML inference pipelines and evaluation workflows.
  • Ability to interpret unstructured model outputs and synthesize guidance.
  • Hands-on experience with LLMs, multimodal models, and NLP systems.
  • Familiarity with AI quality metrics, hallucination detection, and alignment methods.

Responsibilities

  • Lead rigorous model evaluations for LLMs and multimodal models across tasks.
  • Develop evaluation frameworks to quantify human-perceived quality metrics.
  • Translate failure modes into measurable loss patterns and data tweaks.
  • Collaborate with engineering to improve model behavior via prompts and RAG strategies.
  • Map error patterns using ML techniques to guide improvements.
  • Build MLOps workflows and scalable evaluation pipelines for CI/CD.
  • Architect distributed evaluation pipelines for high-throughput testing.

Skills

Python
PyTorch
JAX
Hugging Face
LLMs
Model interpretation
Prompt engineering
RAG architectures

Education

Bachelor’s or Master’s degree in Computer Science / ML / AI or related field

Tools

MLflow
Weigths & Biases
Vector databases
Semantic search
Hugging Face tools

Job description

Summary

We are seeking an AI Evaluation Engineer to help understand, analyse, and improve the performance of advanced AI systems. This role combines data-driven analysis, evaluation design, and human-centered insights to assess AI behavior and drive continuous improvement. You will develop evaluation methodologies, analyse model outputs, identify areas for optimization, and translate findings into actionable recommendations for technical and business stakeholders.

Key Responsibilities
  • Lead Rigorous Model Evaluations: Architect and execute comprehensive evaluation suites for LLMs and multimodal models, identifying edge cases in multi-step reasoning, factuality, adversarial robustness, safety, and alignment.
  • Advanced Scoring Frameworks: Develop deterministic, heuristic, and LLM-assisted evaluation frameworks (e.g., LLM-as-a-judge, reward modeling) to quantify human-perceived quality metrics (e.g., helpfulness, hallucination rates).
  • Actionable Signal Extraction: Translate qualitative failure modes into quantifiable loss patterns, programmatic guardrails, and actionable data-mixture adjustments for model training and inference.
  • Improve Performance: Partner with engineering teams to refine model behavior, leveraging evaluation telemetry to inform prompt engineering, Retrieval-Augmented Generation (RAG) strategies, and model fine-tuning.
  • Latent Pattern Recognition: Apply advanced ML techniques (e.g., embedding-based clustering, representation learning, perturbation analysis) to systematically map error taxonomies and latent failure manifolds in model outputs.
  • MLOps & Automation: Develop robust MLOps workflows to codify evaluation metrics, automate regression testing across model checkpoints, and integrate human-centric assessments into ML CI/CD pipelines.
  • Distributed Evaluation Pipelines: Architect scalable, distributed inference and processing pipelines (e.g., Ray, vLLM) for high-throughput model evaluation, automated annotation, and output analysis at scale.
  • Human-Centric Metrics: Define quantitative evaluation frameworks that capture nuanced human factors, including trust calibration, conversational state tracking, and interpretability.
  • Auto-Evaluator Systems: Build automated evaluation pipelines utilizing LLMs to assess outputs at scale, optimizing for high correlation with human baseline annotations.
  • Cross-Functional Partnership: Collaborate with ML researchers, software developers, and product managers across Apple to translate product requirements into scalable, reliable, and efficient model evaluation infrastructure.
Must Have Skills
  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or a related technical field
  • Relevant industry experience in ML Engineering or Applied Research.
  • Advanced proficiency in Python and modern deep learning ecosystems (PyTorch, JAX, Hugging Face).
  • Proven experience building scalable ML inference pipelines, model-evaluation workflows, and structured rating frameworks for large-scale AI systems.
  • Strong ability to interpret unstructured model outputs (text, transcripts, embedding spaces) and synthesize qualitative findings into actionable engineering guidance and training objectives.
  • Hands-on experience developing, fine-tuning, or evaluating LLMs, multimodal models, and NLP systems.
  • Deep familiarity with AI quality metrics, hallucination detection techniques (e.g., SelfCheckGPT), model alignment (RLHF/DPO), and LLM-as-a-judge frameworks (e.g., G-Eval, DeepEval).
  • Experience building internal tools or automated pipelines for ML workflows using tools like MLflow, Weights & Biases, or similar platforms.
  • Strong familiarity with advanced prompt engineering, RAG architectures (vector databases, semantic search), and Fine-Tuning.
Nice to haves
  • Knowledge of human factors, HCI, or cognitive science methodologies as applied to AI system design.

We regret to inform that only shortlisted candidates will be notified / contacted.

EA Registration number : YAP JIA YI , R25157934

Allegis Group Singapore Pte Ltd, Company Reg No. 200909448N, EA Licence No. 10C4544

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer: Optimize & Validate LLMs
AI Evaluation Engineer: Optimize & Validate LLMs

ALLEGIS GROUP SINGAPORE PRIVATE LIMITED • Singapore

On-site
SGD 120,000 - 160,000
AI Specialist - TEKsystems (Allegis Group Singapore Pte Ltd)
AI Specialist - TEKsystems (Allegis Group Singapore Pte Ltd)

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
ML Evaluation Engineer
ML Evaluation Engineer

ALLEGIS GROUP SINGAPORE PRIVATE LIMITED • Singapore

On-site
SGD 90,000 - 150,000
AI Evaluation Architect
AI Evaluation Architect

Allegis Group Singapore Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
AI Engineer (Evaluation)
AI Engineer (Evaluation)

Nanyang Technological University Singapore • Singapore

On-site
SGD 90,000 - 130,000
Senior AI Evaluation Engineer
Senior AI Evaluation Engineer

Singtel • Singapore

On-site
Confidential
AI Systems Engineer (Model Routing)
AI Systems Engineer (Model Routing)

SignalPlus • Singapore

On-site
SGD 180,000 - 240,000
AI Product Analyst (Developer Community), AI Verify Foundation
AI Product Analyst (Developer Community), AI Verify Foundation

IMDA • Singapore

On-site
SGD 60,000 - 90,000
ML Evaluation Scientist - Data-Driven AI QA
ML Evaluation Scientist - Data-Driven AI QA

ALLEGIS GROUP SINGAPORE PRIVATE LIMITED • Singapore

On-site
SGD 90,000 - 150,000
[T07] AI Engineer - Permanent
[T07] AI Engineer - Permanent

TALENTSIS PTE. LTD. • Singapore

On-site
SGD 150,000 - 190,000