AI Evaluation Engineer — Quality Gatekeeper for AI Agents

Joblogic

Pakistan

Hybrid

PKR 2,500,000 - 4,500,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Professional environment
Market salary
Life Insurance
Medical Insurance (Family)
OPD
Provident Fund
Gym facility
Remote Working
Company trip
29 Annual Leaves
Sick & compassionate leaves
Onsite with UK team

Job summary

Joblogic is building an AI Agent Platform and seeks an AI Evaluation Engineer to own quality for all models integrated into the release. You will define what “good” means, measure it rigorously, and drive fixes across CI and live production traffic.

You will collaborate with engineers, data science, and product to ensure agent reliability for tens of thousands of field-service businesses relying on AI agents acting on their behalf.

Qualifications

  • 3+ years in roles where evaluating models was the core of the job — ML evaluation, ML quality, applied ML, or data science with an evaluation focus.
  • Strong Python engineering skills, with experience building test harnesses and clean, well-tested code.
  • Practical experience evaluating LLM-powered applications or AI agents: building datasets, defining heuristic and LLM-as-judge rubrics, running evaluations, and interpreting results to improve a system.
  • Experience grading tool-using agents on both trajectory and outcome — tool-call correctness, expected-trajectory match, end-state checks.
  • Experience building LLM judges and aligning them to human labels, including agreement statistics and mitigation of bias.

Responsibilities

  • Own release sign-off — gate every agent version on a green regression suite, judge-scored evaluations within agreed error bars, and a red-team pass.
  • Run error analysis — review production traces across chat, email, WhatsApp, and voice; maintain failure taxonomy and create dataset examples.
  • Build datasets and rubrics — design golden datasets and grading rubrics for multi-turn, tool-using agents, promote production runs into datasets.
  • Build and calibrate LLM judges — align judges to human labels, publish agreement per rubric, retire or retrain drifted judges.
  • Run offline and online evaluation — CI regression suites with online scoring and drift tracking.
  • Triage regressions to root cause — attribute failures and propose concrete fixes for engineers.
  • Own voice quality metrics — define gate call-outcome metrics like task success and latency budgets.
  • Evaluate classical ML models — set acceptance criteria, thresholds, and drift alerts.
  • Run adversarial testing — probe for prompt injection, jailbreaks, data leakage, and tool misuse.
  • Run the human evaluation programme — manage annotation queues and inter-annotator agreement.
  • Collaborate & ship — work cross-functionally using Jira and Slack, document clearly, and ship iteratively.

Skills

ML evaluation
Python engineering
LLM evaluation
Data analysis
Cross-functional collaboration

Tools

Jira
LangSmith
PromptFoo
MLflow

Job description

Joblogic is building an AI Agent Platform and seeks an AI Evaluation Engineer to own quality for all models integrated into the release. You will define what “good” means, measure it rigorously, and drive fixes across CI and live production traffic.

You will collaborate with engineers, data science, and product to ensure agent reliability for tens of thousands of field-service businesses relying on AI agents acting on their behalf.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Forward Deployed Engineer - AI Assurance
Forward Deployed Engineer - AI Assurance

Systems Limited • Karachi Division

On-site
PKR 1,800,000 - 4,200,000
AI Assurance Engineer - Quality, Drift & Evaluation
AI Assurance Engineer - Quality, Drift & Evaluation

Systems Limited • Karachi Division

On-site
PKR 1,800,000 - 4,200,000
AI-Driven QA Lead: Automation & Quality Gate Architect
AI-Driven QA Lead: Automation & Quality Gate Architect

KnowledgeCity • Pakistan

On-site
PKR 600,000 - 900,000
Head QA - AI Native
Head QA - AI Native

KnowledgeCity • Pakistan

On-site
PKR 4,000,000 - 8,000,000
Senior AI QA Engineer - SaaS & Automation
Senior AI QA Engineer - SaaS & Automation

qordata • Karachi Division

On-site
PKR 1,800,000 - 3,000,000
QA Lead
QA Lead

KnowledgeCity • Pakistan

On-site
PKR 600,000 - 900,000
AI Production Engineer: Live Workflows & Orchestration
AI Production Engineer: Live Workflows & Orchestration

Smctechsolutions • Haripur

On-site
PKR 1,800,000 - 3,000,000
Senior AI Agent Engineer: Build Autonomous Agents
Senior AI Agent Engineer: Build Autonomous Agents

iGate Technologies • Islamabad

On-site
PKR 2,600,000 - 5,200,000
Fuel allowance
Medical insurance
Free lunch facility (in-house)
+8
AI Evaluation Engineer
AI Evaluation Engineer

Brilliant Systems • Lahore

On-site
PKR 1,800,000 - 3,200,000
AI Quality Analyst: Prompt Design & Personalization Evaluation
AI Quality Analyst: Prompt Design & Personalization Evaluation

Hire Feed • Pakistan

On-site
PKR 800,000 - 1,100,000