Prompt & Evaluation Engineer

Convo

Islamabad

On-site

PKR 1,800,000 - 2,400,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Convo is seeking a Prompt & Evaluation Engineer with 5+ years of experience in applied NLP/LLM engineering to design and optimize prompt chains, agent behaviors, and evaluation frameworks. The role emphasizes translating business requirements into reliable AI solutions through Python-driven evaluation and robust structured-output contracts.

The candidate will develop modular prompt architectures, encode business constraints, and build comprehensive evaluation suites with regression benchmarks,

Qualifications

  • 3+ years in applied NLP or LLM engineering.
  • Experience building systematic evaluation or regression workflows.
  • Hands-on prompt engineering for multi-step, tool-using or structured-output LLM applications.

Responsibilities

  • Design modular prompt chains and structured-output contracts for commercial analysis and advisory workflows.
  • Encode agent constitutions, business constraints and escalation rules defined by the Guild.
  • Create representative evaluation sets, scoring rubrics and regression benchmarks.
  • Implement Python-based evaluation harnesses and automated failure analysis.
  • Version prompts and evaluation assets, track changes and prevent regressions across model updates.
  • Analyze errors with domain experts and convert findings into prompt, data or workflow improvements.

Skills

Prompt engineering
LLM evaluation
Python scripting
Structured outputs
Regression testing
Version control
Business KPI reasoning

Tools

Git

Job description

Job Summary

Convo is looking for a Prompt & Evaluation Engineer with 5+ years of experience in applied NLP/LLM engineering to design and optimize prompt chains, agent behaviors, and evaluation frameworks. The ideal candidate should have strong expertise in prompt engineering, LLM evaluation, Python, structured outputs, regression testing, and failure analysis, with the ability to translate business requirements and commercial use cases into reliable, measurable AI solutions.

Technical mission

Engineer prompt chains, agent constitutions and evaluation suites that convert governed CPG semantics into reliable commercial-agent behavior.

Key Responsibilities
  • Design modular prompt chains and structured-output contracts for commercial analysis and advisory workflows.
  • Encode agent constitutions, business constraints and escalation rules defined by the Guild.
  • Create representative evaluation sets, scoring rubrics and regression benchmarks.
  • Implement Python-based evaluation harnesses and automated failure analysis.
  • Version prompts and evaluation assets, track changes and prevent regressions across model updates.
  • Analyze errors with domain experts and convert findings into prompt, data or workflow improvements.
Required Technical Capabilities
  • Minimum experience: 3+ years in applied NLP or LLM engineering, including at least 1 year building systematic evaluation or regression workflows.
  • Hands-on prompt engineering for multi-step, tool-using or structured-output LLM applications.
  • LLM evaluation design, including rubric-based, deterministic and model-graded approaches.
  • Python scripting for evaluation pipelines and data analysis.
  • Commercial analytics literacy and ability to reason about KPIs, constraints and business decisions.
  • Experience building regression datasets and diagnosing model/prompt failures.
  • Version control and reproducible experiment practices.
Preferred Experience
  • Agent frameworks, tool calls and retrieval-augmented generation.
  • CPG pricing, trade, category or RGM use cases.
  • Evaluation telemetry, prompt registries and experiment-tracking tools.
Expected deliverables / acceptance evidence
  • Versioned prompt-chain and agent-constitution library.
  • Evaluation datasets, rubrics and automated regression suite.
  • Benchmark results and failure taxonomy by commercial use case.
  • Release criteria and change documentation for approved prompt assets.
Primary Interfaces

Works with the CPG Domain Expert, Agent Runtime, Agentic Governance, application teams and AI Test & Dataset Engineering.

What We Have For You
  • Great compensation package, medical benefit for you and your family, free lunch, annual performance-tied increments & performance recognition awards and a great lean and agile work culture!
  • Convo endorses a culture of diversity in all aspects and aims to build a diverse team of amazing individuals!
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Prompt and Evaluation Engineer
Prompt and Evaluation Engineer

Recruitment Intelligence • Islamabad

On-site
PKR 1,800,000 - 3,200,000
Medical benefits
Free lunch
Annual increments
+2
Prompt & Evaluation Engineer — Craft AI Prompts, Metrics & Tests
Prompt & Evaluation Engineer — Craft AI Prompts, Metrics & Tests

Convo • Islamabad

On-site
PKR 1,800,000 - 2,400,000
Lead Prompting & Evaluation Engineer (NLP)
Lead Prompting & Evaluation Engineer (NLP)

Recruitment Intelligence • Islamabad

On-site
PKR 1,800,000 - 3,200,000
Medical benefits
Free lunch
Annual increments
+2
AI Prompt Engineering Intern
AI Prompt Engineering Intern

Aiotac • Islamabad

Hybrid
PKR 167,000 - 234,000
Mentorship
AI / Prompt Engineer
AI / Prompt Engineer

Inspurate • Karachi Division

On-site
PKR 6,967,670 - 11,148,272
Opportunity for rapid career growth
Exposure to modern AI platforms
Collaborate with UK businesses
Forward Deployed Engineer - GenAI
Forward Deployed Engineer - GenAI

Systems Limited • Karachi Division

On-site
PKR 2,000,000 - 3,200,000
Forward Deployed Engineer - GenAI
Forward Deployed Engineer - GenAI

Systems Limited • Islamabad

On-site
PKR 1,800,000 - 3,000,000
Remote AI Prompt Engineer — Build Production Agent Skills
Remote AI Prompt Engineer — Build Production Agent Skills

Pixalate • Lahore

On-site
PKR 19,439,000 - 33,324,000
Casual, Remote Work Environment
Flexible Hours
Monthly Internet Reimbursement
+1
Full-Stack Developer with Claude / AI Experience
Full-Stack Developer with Claude / AI Experience

Perform1 Pvt Ltd. • Karachi Division

On-site
PKR 1,800,000 - 3,200,000
Senior AI Engineer
Senior AI Engineer

Systems Limited • Punjab

On-site
PKR 2,500,000 - 5,000,000