Forward Deployed Engineer - AI Assurance

Systems Limited

Karachi Division

On-site

PKR 1,800,000 - 4,200,000

Full time

23 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Systems Limited in Karachi seeks an experienced QA/test engineer focused on AI-native applications. You will build test plans, automate features, run regression tests across model/prompt/config changes, and red-team AI features for edge cases.

Own the evals framework, define thresholds, integrate eval pipelines into CI/CD, and communicate quality risk to leadership. Requires 5–9 years in QA with 2+ years AI testing, strong Python automation, and familiarity with AI failure modes.

Qualifications

  • 5–9 years QA/test engineering experience, with 2+ years testing AI/ML-powered features.
  • Strong Python-based test automation skills with CI/CD integration.
  • Understands AI-specific failure modes — hallucination, bias, drift, non-determinism.
  • Statistically literate enough to interpret model evaluation metrics, not just pass/fail results.
  • Familiarity with red-teaming methodologies for AI systems.
  • Clear, assertive communicator — willing to block a release over a quality concern.
  • Detail-oriented and methodical under delivery-timeline pressure.
  • Collaborative but independent — doesn't rubber-stamp under delivery pressure.
  • Explains quality risk in business-impact terms, not just technical jargon.
  • Hands-on evals engineering — builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations.
  • Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review.
  • Understands RAG and agent eval metrics — groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency.
  • Experience wiring evals and drift monitoring into CI/CD and production observability.

Responsibilities

  • Build test plans and automation for AI-native application features (functional + AI-specific)
  • Run regression testing across model/prompt/config changes to catch silent quality drift
  • Red-team AI features for edge cases and adversarial inputs where relevant
  • Build automated eval pipelines integrated into CI/CD
  • Partner with AI Architects to define testability requirements before build starts
  • Own the quality gate before any AI feature ships to production
  • Communicate quality risk to delivery leadership in terms they can act on
  • Train delivery teams on AI-specific testing practices
  • Own the evals framework for the practice — golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case
  • Define eval acceptance thresholds per engagement and gate releases on them
  • Build eval engineering tooling — dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read
  • Instrument production evals and drift monitoring, feeding failures back into the golden datasets

Skills

QA/test engineering
AI/ML testing
Python automation
CI/CD integration
Red-teaming for AI
Statistical literacy
Effective communicator
Detail-oriented
Independent worker
Eval frameworks

Tools

OpenAI Evals
Ragas
LangSmith
Azure AI Foundry evaluations

Job description

Owns quality for AI-native applications — functional testing plus the AI-specific evaluation (accuracy, drift, hallucination).

KEY RESPONSIBILITIES
  • Build test plans and automation for AI-native application features (functional + AI-specific)
  • Run regression testing across model/prompt/config changes to catch silent quality drift
  • Red-team AI features for edge cases and adversarial inputs where relevant
  • Build automated eval pipelines integrated into CI/CD
  • Partner with AI Architects to define testability requirements before build starts
  • Own the quality gate before any AI feature ships to production
  • Communicate quality risk to delivery leadership in terms they can act on
  • Train delivery teams on AI-specific testing practices
  • Own the evals framework for the practice — golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case
  • Define eval acceptance thresholds per engagement and gate releases on them
  • Build eval engineering tooling — dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read
  • Instrument production evals and drift monitoring, feeding failures back into the golden datasets
REQUIREMENTS & SKILLS
  • 5–9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specifically
  • Strong test automation skills (Python-based frameworks, CI/CD integration)
  • Understands AI-specific failure modes — hallucination, bias, drift, non-determinism — and designs tests for them
  • Statistically literate enough to interpret model evaluation metrics, not just pass/fail results
  • Familiarity with red-teaming methodologies for AI systems
  • Clear, assertive communicator — willing to block a release over a quality concern
  • Detail-oriented and methodical under delivery-timeline pressure
  • Collaborative but independent — doesn't rubber-stamp under delivery pressure
  • Explains quality risk in business-impact terms, not just technical jargon
  • Hands-on evals engineering — builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations
  • Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review
  • Understands RAG and agent eval metrics — groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency
  • Experience wiring evals and drift monitoring into CI/CD and production observability
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Forward Deployed Engineer - AI Assurance
Forward Deployed Engineer - AI Assurance

Systems Limited • Islamabad

On-site
PKR 150,000 - 210,000
Forward Deployed Engineer - AI Assurance
Forward Deployed Engineer - AI Assurance

Systems Limited • Lahore

On-site
PKR 2,400,000 - 3,600,000
AI Quality Assurance & Evaluation Engineer
AI Quality Assurance & Evaluation Engineer

Systems Limited • Islamabad

On-site
PKR 150,000 - 210,000
AI Quality & Evaluation Engineer
AI Quality & Evaluation Engineer

Systems Limited • Lahore

On-site
PKR 2,400,000 - 3,600,000
Head QA - AI Native
Head QA - AI Native

KnowledgeCity • Pakistan

On-site
PKR 4,000,000 - 8,000,000
AI Architect
AI Architect

Creativechaos • Pakistan

On-site
Forward Deployed Engineer - Digital
Forward Deployed Engineer - Digital

Systems Limited • Karachi Division

On-site
PKR 1,800,000 - 3,200,000
AI Assurance Engineer - Quality, Drift & Evaluation
AI Assurance Engineer - Quality, Drift & Evaluation

Systems Limited • Karachi Division

On-site
PKR 1,800,000 - 4,200,000
Forward Deployed Engineer - Digital
Forward Deployed Engineer - Digital

Systems Limited • Lahore

On-site
PKR 1,200,000 - 3,000,000
AI Governance
AI Governance

Systems Limited • Lahore

On-site
PKR 3,500,000 - 7,000,000