Forward Deployed Engineer - AI Assurance

Systems Limited

Malaysia

On-site

MYR 60,000 - 120,000

Full time

38 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Systems Limited is seeking a QA/test engineer to build and maintain automated tests for AI-native application features, ensuring functional quality and AI-specific evaluation metrics. The role focuses on regression testing across model/prompt/config changes, red-teaming for edge cases, and integrating eval pipelines into CI/CD.

The candidate will own the quality gates for AI features, communicate risks to leadership, and train delivery teams in AI-focused testing practices.

Qualifications

  • 5–9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specifically
  • Strong test automation skills (Python-based frameworks, CI/CD integration)
  • Understands AI-specific failure modes — hallucination, bias, drift, non-determinism
  • Statistically literate enough to interpret model evaluation metrics, not just pass/fail results
  • Familiarity with red-teaming methodologies for AI systems
  • Clear, assertive communicator — willing to block a release over a quality concern
  • Detail-oriented and methodical under delivery-timeline pressure
  • Collaborative but independent — doesn't rubber-stamp under delivery pressure
  • Explains quality risk in business-impact terms, not just technical jargon
  • Hands-on evals engineering — builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations
  • Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review
  • Understands RAG and agent eval metrics — groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency
  • Experience wiring evals and drift monitoring into CI/CD and production observability

Responsibilities

  • Build test plans and automation for AI-native application features (functional + AI-specific)
  • Run regression testing across model/prompt/config changes to catch silent quality drift
  • Red-team AI features for edge cases and adversarial inputs where relevant
  • Build automated eval pipelines integrated into CI/CD
  • Partner with AI Architects to define testability requirements before build starts
  • Own the quality gate before any AI feature ships to production
  • Communicate quality risk to delivery leadership in terms they can act on
  • Train delivery teams on AI-specific testing practices
  • Own the evals framework for the practice — golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case
  • Define eval acceptance thresholds per engagement and gate releases on them
  • Build eval engineering tooling — dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read
  • Instrument production evals and drift monitoring, feeding failures back into the golden datasets

Skills

QA/testing expertise
Test automation
Python automation
CI/CD integration
AI/ML testing
Statistical thinking
Red-teaming methods
Communication skills
Attention to detail

Tools

OpenAI Evals
Ragas
DeepEval
LangSmith
Azure AI Foundry

Job description

Owns quality for AI-native applications — functional testing plus the AI-specific evaluation (accuracy, drift, hallucination).

KEY RESPONSIBILITIES
  • Build test plans and automation for AI-native application features (functional + AI-specific)
  • Run regression testing across model/prompt/config changes to catch silent quality drift
  • Red-team AI features for edge cases and adversarial inputs where relevant
  • Build automated eval pipelines integrated into CI/CD
  • Partner with AI Architects to define testability requirements before build starts
  • Own the quality gate before any AI feature ships to production
  • Communicate quality risk to delivery leadership in terms they can act on
  • Train delivery teams on AI-specific testing practices
  • Own the evals framework for the practice — golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case
  • Define eval acceptance thresholds per engagement and gate releases on them
  • Build eval engineering tooling — dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read
  • Instrument production evals and drift monitoring, feeding failures back into the golden datasets
REQUIREMENTS & SKILLS
  • 5–9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specifically
  • Strong test automation skills (Python-based frameworks, CI/CD integration)
  • Understands AI-specific failure modes — hallucination, bias, drift, non-determinism — and designs tests for them
  • Statistically literate enough to interpret model evaluation metrics, not just pass/fail results
  • Familiarity with red-teaming methodologies for AI systems
  • Clear, assertive communicator — willing to block a release over a quality concern
  • Detail-oriented and methodical under delivery-timeline pressure
  • Collaborative but independent — doesn't rubber-stamp under delivery pressure
  • Explains quality risk in business-impact terms, not just technical jargon
  • Hands-on evals engineering — builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations
  • Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review
  • Understands RAG and agent eval metrics — groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency
  • Experience wiring evals and drift monitoring into CI/CD and production observability
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Forward-Deployed AI QA & Eval Engineer
Forward-Deployed AI QA & Eval Engineer

Systems Limited • Malaysia

On-site
MYR 60,000 - 120,000
AI Engineer (Test Engineering)
AI Engineer (Test Engineering)

MetLife • Kuala Lumpur

On-site
MYR 90,000 - 150,000
AI Governance
AI Governance

Systems Limited • Malaysia

On-site
MYR 180,000 - 320,000
Forward Deployed Engineer - GenAI
Forward Deployed Engineer - GenAI

Systems Limited • Malaysia

On-site
MYR 120,000 - 180,000
AI Deployment Strategist
AI Deployment Strategist

Systems Limited • Kuala Lumpur

On-site
MYR 150,000 - 240,000
QA Azure AI Engineer EDV WERKE AG Building on experience, results and commitment.
QA Azure AI Engineer EDV WERKE AG Building on experience, results and commitment.

Edvwerke • Kuala Lumpur

On-site
MYR 100,000 - 130,000
Competitive salary with performance-based bonuses
Opportunities for professional development
Dynamic and collaborative work environment
Senior AI Solutions Engineer (Team Lead)
Senior AI Solutions Engineer (Team Lead)

techstreet • Petaling Jaya

On-site
MYR 180,000 - 280,000
AI Software Engineer
AI Software Engineer

Webby Group • Kuala Lumpur

On-site
MYR 120,000 - 160,000
Senior AI Solutions Engineer (Team Lead)
Senior AI Solutions Engineer (Team Lead)

techstreet • Selangor

On-site
MYR 180,000 - 340,000
AI DevOps Engineer (MLOps & Cloud)
AI DevOps Engineer (MLOps & Cloud)

DXC Technology Inc. • Petaling Jaya

On-site
MYR 100,000 - 150,000