Forward Deployed Engineer - AI Assurance

Systems Limited

Lahore

On-site

PKR 2,400,000 - 3,600,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Systems Limited is seeking a QA/test engineer to lead automation for AI-native features. You will design evaluation harnesses and run regression tests to catch quality drift across model changes, with emphasis on accuracy, hallucination, and bias.

You will collaborate with AI architects, own the eval framework, and communicate quality risk to leadership, driving gated releases and robust test coverage in a fast-paced delivery environment.

Qualifications

  • 5–9 yrs QA/test engineering with AI/ML feature testing.
  • Strong Python-based test automation and CI/CD integration.
  • Familiarity with AI-specific failure modes and evaluation metrics.
  • Ability to interpret model evaluation metrics beyond pass/fail.

Responsibilities

  • Build test plans and automation for AI-native features (functional + AI-specific).
  • Design evaluation harnesses for model outputs—accuracy, drift, hallucination rate.
  • Run regression tests across model/prompt/config changes to detect quality drift.
  • Red-team AI features for edge cases and adversarial inputs.
  • Create CI/CD integrated eval pipelines and dashboards for teams.
  • Define eval thresholds and gate releases on them, with clear risk communication.
  • Own the evals framework including datasets, rubrics, and LLm calibration.

Skills

QA automation
Python
CI/CD
AI/ML testing
Statistical literacy
Red-teaming
Communication
Drift testing

Tools

OpenAI Evals
Ragas
DeepEval
LangSmith
Azure AI Foundry

Job description

ABOUT:

Owns quality for AI-native applications - functional testing plus the AI-specific evaluation (accuracy, drift, hallucination).

KEY RESPONSIBILITIES
  • Build test plans and automation for AI-native application features (functional + AI-specific)
  • Design evaluation harnesses for model/agent outputs - accuracy, consistency, hallucination rate
  • Run regression testing across model/prompt/config changes to catch silent quality drift
  • Red-team AI features for edge cases and adversarial inputs where relevant
  • Build automated eval pipelines integrated into CI/CD
  • Partner with AI Architects to define testability requirements before build starts
  • Own the quality gate before any AI feature ships to production
  • Communicate quality risk to delivery leadership in terms they can act on
  • Train delivery teams on AI-specific testing practices
  • Own the evals framework for the practice - golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case
  • Define eval acceptance thresholds per engagement and gate releases on them
  • Build eval engineering tooling - dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read
  • Instrument production evals and drift monitoring, feeding failures back into the golden datasets
REQUIREMENTS & SKILLS
  • 5-9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specifically
  • Strong test automation skills (Python-based frameworks, CI/CD integration)
  • Understands AI-specific failure modes - hallucination, bias, drift, non-determinism - and designs tests for them
  • Statistically literate enough to interpret model evaluation metrics, not just pass/fail results
  • Familiarity with red-teaming methodologies for AI systems
  • Clear, assertive communicator - willing to block a release over a quality concern
  • Detail-oriented and methodical under delivery-timeline pressure
  • Collaborative but independent - doesn't rubber-stamp under delivery pressure
  • Explains quality risk in business-impact terms, not just technical jargon
  • Hands-on evals engineering - builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations
  • Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review
  • Understands RAG and agent eval metrics - groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency
  • Experience wiring evals and drift monitoring into CI/CD and production observability
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Forward Deployed Engineer - AI Assurance
Forward Deployed Engineer - AI Assurance

Systems Limited • Islamabad

On-site
PKR 150,000 - 210,000
AI Quality Assurance & Evaluation Engineer
AI Quality Assurance & Evaluation Engineer

Systems Limited • Islamabad

On-site
PKR 150,000 - 210,000
AI Quality & Evaluation Engineer
AI Quality & Evaluation Engineer

Systems Limited • Lahore

On-site
PKR 2,400,000 - 3,600,000
AI Governance
AI Governance

Systems Limited • Lahore

On-site
PKR 3,500,000 - 7,000,000
AI Governance
AI Governance

Systems Limited • Islamabad

On-site
PKR 1,000,000 - 2,000,000
AI Architect
AI Architect

Creativechaos • Pakistan

On-site
QA Lead - AI Native
QA Lead - AI Native

KnowledgeCity • Pakistan

On-site
PKR 2,500,000 - 4,200,000
Forward Deployed Engineer - GenAI
Forward Deployed Engineer - GenAI

Systems Limited • Lahore

On-site
PKR 2,500,000 - 5,500,000
Director of Innovation Delivery
Director of Innovation Delivery

Creativechaos • Pakistan

Remote
PKR 2,000,000 - 3,000,000
Senior Quality Assurance Engineer
Senior Quality Assurance Engineer

Integriti Global • Lahore

On-site
PKR 3,000,000 - 6,000,000