AI Evaluation Engineer

Joblogic Service Management Software

Lahore

On-site

PKR 900,000 - 1,500,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Life Insurance
Medical Insurance (Family)
Provident Fund
Gym Facility
Remote Working (During Pandemic)
Company trip
29 Annual Leaves
Sick & uncapped Compassionate Leaves
Onsite with UK team

Job summary

Joblogic is building its AI Agent Platform and seeks an AI Evaluation Engineer to own model quality across channels and tools. You will define rubrics, design datasets, and implement release gates that guard every agent version in CI and live traffic.

You will collaborate with engineers, data, and product to ensure robust evaluations, actionable insights, and continuous improvement of agent performance in production. This is a hybrid role based in Lahore, Pakistan.

Qualifications

  • 3+ years evaluating ML models.
  • Strong Python engineering skills.
  • Experience evaluating LLM-powered applications or AI agents.
  • Experience grading tool-using agents and reporting reliability.
  • Knowledge of ML evaluation metrics and drift monitoring.
  • Data analysis using Pandas/NumPy/SQL.
  • Awareness of AI safety and GDPR implications.
  • Cross-functional collaboration with data/product teams.

Responsibilities

  • Own release sign-off with a green regression suite and judge-scored evaluations.
  • Run error analysis on production traces across channels.
  • Build datasets and rubrics for multi-turn, tool-using agents.
  • Build and calibrate LLM judges aligned to human labels.
  • Run offline and online evaluation with CI and production traffic.
  • Triage regressions to root cause and propose fixes.
  • Define voice quality metrics and latency budgets.
  • Evaluate ML models across vision/speech/predictive domains.
  • Conduct adversarial testing and track regression coverage.
  • Lead human-evaluation programs with clear guidelines.

Skills

Python
ML evaluation
LLM evaluation
Agent evaluation
LLM judges
ML fundamentals
Pandas
NumPy
SQL
Prompt engineering
Jira
Slack

Tools

LangSmith
LangGraph
MLflow
PromptFoo

Job description

Job Description:
The Joblogic Story

Established in 1998, Joblogic is the UK’s #1 Field Service Management (FSM) software platform. We are a global business with offices in the UK, Pakistan, and Vietnam. Since our management buy-out in 2013, we have grown from ~£500K ARR to ~£35M+ ARR and expanded our team from 11 to 500+ people.

Recently, we secured a strategic growth investment from Vista Equity Partners - a global technology investor specialising in enterprise software. This investment includes over £100 million in new primary capital and will fuel our next phase of growth by accelerating our AI-first roadmap, expanding our platform into CAFM (Computer-Aided Facilities Management) capabilities, and supporting our expansion across Europe and beyond.

With Vista’s backing, we’re transforming from a successful UK business into a global scaling SaaS rocket ship - and we’d love for you to join us on our journey to £100M ARR across international markets.

Joblogic provides software to service contractors who install and maintain the built environment. Our platform helps businesses streamline operations, improve profitability, ensure compliance, and achieve rapid growth. With over 100,000 users across industries including HVAC, plumbing, electrical maintenance, facilities management, and building fabric maintenance, we are entering a new era of intelligent automation, predictive maintenance, and data-driven decision‑making for service firms.

About the Role

We are building Joblogic’s AI Agent Platform - a multi-tenant system for designing, versioning, evaluating, and running AI agents that work across email, voice, SMS, WhatsApp, and CRM channels on behalf of our customers. The platform is built on a LangGraph runtime with retrieval over Azure AI Search, a real-time voice stack, human-in-the-loop review queues, and an evaluation harness backed by LangSmith and PromptFoo.

We are looking for an AI Evaluation Engineer to own quality for everything we ship that has a model in it. You will define what “good” means, measure it rigorously, find out why it is not met, and drive the fixes. The role is deliberately hybrid: you design the rubrics, datasets, and judges, and you build the harnesses and release gates that run them in CI and against live production traffic. Evaluation is not a reporting function here - it is the mechanism by which agent quality improves release over release, and you own it.

You will work closely with the engineers building agents, the data team, and product, and your work will directly determine what tens of thousands of field-service businesses experience when an agent answers on their behalf.

What You’ll Do
  • Own release sign-off - gate every agent version on a green regression suite, judge-scored evaluations within agreed error bars, and a red-team pass - no promotion without them.
  • Run error analysis - review sampled production traces across chat, email, WhatsApp, and voice every week, maintain the failure taxonomy, and turn new failure modes into dataset examples within the sprint.
  • Build datasets and rubrics - design and maintain golden datasets and grading rubrics for multi-turn, tool-using agents, promoting interesting production runs into datasets in LangSmith.
  • Build and calibrate LLM judges - align judges to human labels, re-label a held-out set each cycle, publish agreement per rubric, and retire or retrain judges that drift.
  • Run offline and online evaluation - regression suites in CI alongside online scoring of sampled production traffic, with clear pass/fail thresholds and drift tracking.
  • Triage regressions to root cause - attribute failures to prompt, tool, retrieval, model upgrade, or speech provider using paired statistics rather than aggregate deltas, and hand engineers concrete fixes.
  • Own voice quality metrics - define and gate call-outcome metrics - task success, containment, barge-in recovery, word error rate under noise - alongside latency budgets.
  • Evaluate classical ML models - set acceptance criteria, slice-level thresholds, calibration checks, and drift alerts for the vision, speech, and predictive models the platform depends on.
  • Run adversarial testing - probe for prompt injection through inbound email and messaging content, jailbreaks, PII leakage, and tool misuse, and add regression coverage for every finding.
  • Run the human evaluation programme - own annotation queues, reviewer guidelines, and inter-annotator agreement as an ongoing operation rather than a one-off study, and keep human and automated scores connected.
  • Collaborate & ship - work in a cross-functional team using tools such as Jira and Slack, write clear documentation, and ship iteratively with a strong quality bar.
Essential Experience and Skills
  • 3+ years in roles where evaluating models was the core of the job — ML evaluation, ML quality, applied ML, or data science with an evaluation focus.
  • Strong Python engineering skills, with experience building test harnesses and clean, well-tested code.
  • Practical experience evaluating LLM-powered applications or AI agents: building datasets, defining heuristic and LLM-as-judge rubrics, running evaluations, and interpreting results to improve a system. This is a core requirement.
  • Experience grading tool-using agents on both trajectory and outcome — tool-call correctness, expected-trajectory match, end-state checks — and reporting reliability over repeated trials.
  • Experience building LLM judges and aligning them to human labels, including agreement statistics and mitigation of position, verbosity, and self-preference bias.
  • Solid classical ML evaluation foundations: classification, regression, and ranking metrics, cross-validation, calibration, slice-based evaluation, and drift monitoring.
  • Statistics for small evaluation sets: paired comparisons, confidence intervals, power analysis, and minimum detectable effect — you know roughly how many examples a claim needs before you make it.
  • Experience with RAG evaluation: faithfulness, groundedness, context precision and recall.
  • Hands-on experience with an LLM observability and evaluation platform (LangSmith, MLflow, or equivalent): datasets, experiments, custom evaluators, feedback, and wiring evaluations into release gates.
  • Strong data analysis skills using Pandas, NumPy, and SQL to quantify behaviour and communicate findings.
  • Working knowledge of how agents are built — prompts, tools, retrieval, and memory — and how they interact, sufficient to root-cause a failure rather than only report it.
  • Awareness of AI safety and adversarial risk: prompt injection, jailbreaks, data leakage, tool misuse, and responsible-AI practice.
  • Awareness of the compliance side of evaluation: handling conversation data under UK GDPR, keeping evaluation evidence auditable, and emerging record‑keeping expectations for AI systems.
  • Committed to continuous learning, proactive problem-solving, and timely issue identification, with a keen interest in staying current with a fast-moving field.
  • Strong communicator, experienced in collaborating with cross-functional teams using tools such as Jira and Slack.
  • Creative and innovative thinker, consistently contributing fresh ideas and solutions in alignment with current technological trends.
Nice to Have
  • Experience with LangGraph and LangSmith specifically (tracing, datasets, online evaluators, judge alignment).
  • Experience with evaluation tooling such as PromptFoo, DeepEval, RAGAS, or Inspect.
  • Experience with red-team tooling such as PromptFoo red team, Microsoft PyRIT, or NVIDIA Garak, and familiarity with the OWASP Top 10 for LLM and Agentic Applications.
  • Experience evaluating voice agents: call simulation, turn-taking and latency budgets, speech recognition accuracy under noise.
  • Experience evaluating computer-vision or speech models (mAP/IoU, WER/CER) with sample-level failure mining.
  • Experience with Databricks or AWS SageMaker and experiment tracking with MLflow.
  • Experience designing human-in-the-loop annotation programmes and measuring inter-annotator agreement.
  • Experience with Python web frameworks (FastAPI / Flask), pytest-native harnesses, and CI/CD.
  • Publications, open-source contributions, or public evaluation work.
What We Offer
  • Professional Working environment
  • Market Competitive Salary
  • Life Insurance & Medical Insurance (Including Family)
  • OPD
  • Provident Fund
  • Gym Facility
  • Maximum 45 Weekly Hours (Monday-Friday)
  • Remote Working (During Pandemic Situation)
  • Company trip
  • 29 Annual Leaves
  • 8 Sick & uncapped Compassionate Leaves (As per Company Policy)
  • Have a chance to work onsite with the UK team

The interview will be held in multiple stages where the successful candidate must demonstrate they meet the essential requirement criteria and have the relevant experience. You must be able to travel to our Lahore office daily.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

Joblogic Service Management Software • Lahore

On-site
PKR 3,000,000 - 5,400,000
Professional Working environment
Market Competitive Salary
Life Insurance & Medical Insurance (In
Remote AI Evaluation Engineer
Remote AI Evaluation Engineer

Joblogic Service Management Software • Lahore

On-site
PKR 900,000 - 1,500,000
Life Insurance
Medical Insurance (Family)
Provident Fund
+6
AI Product Engineer
AI Product Engineer

Exceptional Dental • Karachi Division

Hybrid
PKR 33,620,000 - 48,562,000
Private medical insurance
Paid holiday
Career development opportunities
Senior Full Stack AI Engineer
Senior Full Stack AI Engineer

PRG Pakistan • Lahore

On-site
PKR 2,500,000 - 3,500,000
Engineering Lead
Engineering Lead

Travel Innovation Group • Lahore

On-site
Medical Insurance
Pension Scheme
Gym Passport
+3
AI / LLM Engineer
AI / LLM Engineer

HRBS Global • Islamabad

On-site
PKR 1,800,000 - 3,200,000
Senior Software Engineer (AI/ML)
Senior Software Engineer (AI/ML)

Devsinc, LLC • Islamabad

On-site
Provident Fund
Medical Inpatient Facility
Medical Outpatient Facility
+7
AI Developer — Onsite, Sialkot Office
AI Developer — Onsite, Sialkot Office

FabTechSol • Sialkot

On-site
AI Engineer — Agents & RAG
AI Engineer — Agents & RAG

DevMations • Pakistan

On-site
PKR 1,800,000 - 3,000,000
Senior .Net Engineer
Senior .Net Engineer

VIDIZMO LLC • Karachi Division

On-site
PKR 2,400,000 - 4,800,000
Health Insurance
Maternity Cover
Leave Encashment
+7