AI Evaluation Engineer

Brilliant Systems

Lahore

On-site

PKR 1,800,000 - 3,200,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Brilliant Systems is seeking an engineer to own the evaluation of ML/LLM systems, turning client requirements into measurable evaluation sets and robust tooling. You will build harnesses, monitor production drift, and ensure evaluation is integral to the product lifecycle.

You must demonstrate statistical literacy, strong Python skills, and a knack for developer-friendly tooling, while candidly reporting results that may be inconvenient.

Qualifications

  • You have built evaluation for a machine learning or LLM system that others relied on.
  • You understand what small samples can and cannot tell you.
  • You write code and build tooling that engineers actually adopt.
  • You are willing to report results that nobody wants to hear.

Responsibilities

  • Turn client requirements into measurable evaluation sets and adversarial cases.
  • Build evaluation harnesses and tooling so engineers can run and read them.
  • Track drift in production and raise it early before clients detect it.
  • Contribute to BotUp where evaluation is part of the product, not just internal tooling.
  • Push back early on unmeasurable scope while changes are inexpensive.

Skills

Strong Python
Statistical literacy
Taste for tooling
Scepticism

Tools

Python

Job description

Decide what done means. If we cannot score it, we have not specified it, and we certainly cannot sell it.

We sell fixed-price phases, so we need a defensible answer to whether a thing is finished. For deterministic software that is a test suite. For a system that is probabilistic by design it is an evaluation harness, and building good ones is a speciality. This role owns that speciality across client engagements and our own products.

What you will do
  • Turn client requirements into measurable evaluation sets, including the adversarial cases nobody asked for.
  • Build the harnesses, and the tooling around them, so any engineer can run and read them.
  • Track drift in production and raise it before a client does.
  • Work on BotUp, where every worker is reviewed and every run is metered and logged, so evaluation is part of the product rather than internal tooling.
  • Push back early on scope that cannot be measured, while it is still cheap to change.
What we need from you
  • You have built evaluation for a machine learning or LLM system that other people then relied on.
  • Statistical literacy: you know what a small sample can and cannot tell you.
  • Strong Python, and a taste for tooling other engineers actually adopt.
  • Scepticism. This role only works if you are willing to report a result nobody wants.
Useful but not required
  • A testing or QA engineering background before moving into machine learning.
  • Experience with human-in-the-loop annotation at any scale.
  • You have written about evaluation somewhere public.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Evaluation Architect & Quality Harness Engineer
ML Evaluation Architect & Quality Harness Engineer

Brilliant Systems • Lahore

On-site
PKR 1,800,000 - 3,200,000
Forward Deployed Engineer - AI Assurance
Forward Deployed Engineer - AI Assurance

Systems Limited • Karachi Division

On-site
PKR 1,800,000 - 4,200,000
Senior AI Engineer
Senior AI Engineer

Brilliant Systems • Lahore

On-site
PKR 1,800,000 - 2,400,000
MLOps Engineer
MLOps Engineer

Joblogic • Pakistan

Hybrid
PKR 2,500,000 - 4,500,000
Professional environment
Market salary
Life Insurance
+9
Machine Learning Engineer, Retrieval
Machine Learning Engineer, Retrieval

Brilliant Systems • Lahore

On-site
PKR 600,000 - 1,200,000
Prompt & Evaluation Engineer — Craft AI Prompts, Metrics & Tests
Prompt & Evaluation Engineer — Craft AI Prompts, Metrics & Tests

Convo • Islamabad

On-site
PKR 1,800,000 - 2,400,000
Prompt & Evaluation Engineer
Prompt & Evaluation Engineer

Convo • Islamabad

On-site
PKR 1,800,000 - 2,400,000
AI Evaluation Engineer — Quality Gatekeeper for AI Agents
AI Evaluation Engineer — Quality Gatekeeper for AI Agents

Joblogic • Pakistan

Hybrid
PKR 2,500,000 - 4,500,000
Professional environment
Market salary
Life Insurance
+9
AI Engineer
AI Engineer

Tkxel LLC • Lahore

On-site
PKR 38,770,000 - 52,617,000
Senior NLP & Machine Learning Engineer
Senior NLP & Machine Learning Engineer

LimeoX LLC • Sargodha

On-site
PKR 3,080,000 - 4,620,000