Forward Deployed Engineer - AI Assurance

Systems Limited

Islamabad

On-site

PKR 150,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Systems Limited is seeking a QA/test engineer focused on AI-native applications. You will design and build test plans and automation for AI features, and run evaluations of model outputs within CI/CD pipelines.

You will own the evals framework, including golden datasets, rubrics, and monitoring for drift and hallucination, and you will communicate risks to delivery leadership while training teams on best practices.

Qualifications

  • 5-9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specifically.
  • Strong test automation skills (Python-based frameworks, CI/CD integration).
  • Understands AI-specific failure modes - hallucination, bias, drift, non-determinism - and designs tests for them
  • Statistically literate enough to interpret model evaluation metrics, not just pass/fail results
  • Familiarity with red-teaming methodologies for AI systems
  • Clear, assertive communicator - willing to block a release over a quality concern

Responsibilities

  • Build test plans and automation for AI-native application features (functional + AI-specific)
  • Design evaluation harnesses for model/agent outputs - accuracy, consistency, hallucination rate
  • Run regression testing across model/prompt/config changes to catch silent quality drift
  • Red-team AI features for edge cases and adversarial inputs where relevant
  • Build automated eval pipelines integrated into CI/CD
  • Partner with AI Architects to define testability requirements before build starts
  • Own the quality gate before any AI feature ships to production
  • Communicate quality risk to delivery leadership in terms they can act on
  • Train delivery teams on AI-specific testing practices
  • Own the evals framework for the practice - golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case
  • Define eval acceptance thresholds per engagement and gate releases on them
  • Build eval engineering tooling - dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read
  • Instrument production evals and drift monitoring, feeding failures back into the golden datasets

Skills

QA Automation
Python
CI/CD
AI/ML Testing
Statistical Analysis
Red-Teaming
Clear Communication
Deterministic Testing

Tools

OpenAI Evals
Ragas
DeepEval
LangSmith
Azure AI Foundry

Job description

ABOUT:

Owns quality for AI-native applications - functional testing plus the AI-specific evaluation (accuracy, drift, hallucination).

KEY RESPONSIBILITIES
  • Build test plans and automation for AI-native application features (functional + AI-specific)
  • Design evaluation harnesses for model/agent outputs - accuracy, consistency, hallucination rate
  • Run regression testing across model/prompt/config changes to catch silent quality drift
  • Red-team AI features for edge cases and adversarial inputs where relevant
  • Build automated eval pipelines integrated into CI/CD
  • Partner with AI Architects to define testability requirements before build starts
  • Own the quality gate before any AI feature ships to production
  • Communicate quality risk to delivery leadership in terms they can act on
  • Train delivery teams on AI-specific testing practices
  • Own the evals framework for the practice - golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case
  • Define eval acceptance thresholds per engagement and gate releases on them
  • Build eval engineering tooling - dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read
  • Instrument production evals and drift monitoring, feeding failures back into the golden datasets
REQUIREMENTS & SKILLS
  • 5-9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specifically
  • Strong test automation skills (Python-based frameworks, CI/CD integration)
  • Understands AI-specific failure modes - hallucination, bias, drift, non-determinism - and designs tests for them
  • Statistically literate enough to interpret model evaluation metrics, not just pass/fail results
  • Familiarity with red-teaming methodologies for AI systems
  • Clear, assertive communicator - willing to block a release over a quality concern
  • Detail-oriented and methodical under delivery-timeline pressure
  • Collaborative but independent - doesn't rubber-stamp under delivery pressure
  • Explains quality risk in business-impact terms, not just technical jargon
  • Hands-on evals engineering - builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations
  • Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review
  • Understands RAG and agent eval metrics - groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency
  • Experience wiring evals and drift monitoring into CI/CD and production observability
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Forward Deployed Engineer - AI Assurance
Forward Deployed Engineer - AI Assurance

Systems Limited • Lahore

On-site
PKR 2,400,000 - 3,600,000
AI Quality Assurance & Evaluation Engineer
AI Quality Assurance & Evaluation Engineer

Systems Limited • Islamabad

On-site
PKR 150,000 - 210,000
AI Quality & Evaluation Engineer
AI Quality & Evaluation Engineer

Systems Limited • Lahore

On-site
PKR 2,400,000 - 3,600,000
AI Governance
AI Governance

Systems Limited • Lahore

On-site
PKR 3,500,000 - 7,000,000
AI Governance
AI Governance

Systems Limited • Islamabad

On-site
PKR 1,000,000 - 2,000,000
AI Architect
AI Architect

Creativechaos • Pakistan

On-site
QA Lead - AI Native
QA Lead - AI Native

KnowledgeCity • Pakistan

On-site
PKR 2,500,000 - 4,200,000
Forward Deployed Engineer - GenAI
Forward Deployed Engineer - GenAI

Systems Limited • Lahore

On-site
PKR 2,500,000 - 5,500,000
Director of Innovation Delivery
Director of Innovation Delivery

Creativechaos • Pakistan

Remote
PKR 2,000,000 - 3,000,000
Senior Quality Assurance Engineer
Senior Quality Assurance Engineer

Integriti Global • Lahore

On-site
PKR 3,000,000 - 6,000,000