67503 Agent Evaluation & Instrumentation Engineer

Cephas Consultancy Services Private Limited

Pune District

On-site

INR 3,000,000 - 4,500,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cephas Consultancy Services Private Limited in Pune is seeking an Agent Evaluation & Instrumentation Engineer with 8-11 years experience to lead the design, calibration, and adjudication of evaluation suites and gate thresholds across agent archetypes.

You will own the regression-pack methodology for AI/ML model changes, run gate reviews, calibrate thresholds to production reality, and mentor team members while delivering monthly quality reports on eval trends and defect clusters.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related field
  • 6+ years in ML/data/software with strong evaluation or quality focus
  • Hands-on experience evaluating LLM or ML systems
  • Experience with agentic systems or RAG pipelines
  • At least 3技能: LLM evaluation design, Python tooling, data analysis
  • Analytical rigour and attention to detail
  • Clear technical writing for gate findings
  • Ability to influence build teams on quality and present to management

Responsibilities

  • Design and maintain offline evaluation suites (golden sets, regression packs, adversarial/safety probes) and the continuous-evaluation scoring pipeline across archetypes.
  • Conduct operability gate reviews: assess evidence packs, reproduce evaluation results, and recommend go/no-go with documented findings.
  • Own the model-update regression pack methodology and adjudicate regression runs against archived baselines with AIML Operations team.
  • Calibrate gate thresholds against production reality and maintain evaluation drift hygiene.
  • Produce the monthly quality report per agent: eval trends, failure-mode taxonomy, and defect clusters with reproduction traces.
  • Mentor members of the team when needed and review their work for quality and consistency.

Skills

LLM evaluation design
Python tooling
Data analysis
Observability tools
Adversarial testing
Technical writing
Management presentation
Stakeholder influence

Education

Bachelor’s or Master’s degree in Computer Science

Tools

promptfoo
DeepEval
Custom evaluation harness

Job description

67503 Agent Evaluation & Instrumentation Engineer
About this position

Positions:1 Full Time
Experience
8 - 11 Years

SECTION A: POSITION SUMMARY

State the objective and purpose of the role.

Serve as the technical evaluator within the team, responsible for the design, calibration, and adjudication of evaluation suites and gate thresholds across agent archetypes. Lead gate reviews under the Team Lead’s authority, own the regression-pack methodology for AI/ML model changes, and act as the technical custodian of evaluation quality and drift hygiene.

SECTION B: KEY RESPONSIBILITIES AND RESULTS

Indicate key responsibilities and performance indicators of this role.

For existing role, please indicate additional responsibilities in bold.

  1. Design and maintain offline evaluation suites (golden sets, regression packs, adversarial/safety probes) and the continuous-evaluation scoring pipeline across archetypes.
  2. Conduct operability gate reviews: assess evidence packs, reproduce evaluation results, and recommend go/no-go with documented findings.
  3. Own the model-update regression pack methodology and adjudicate regression runs against archived baselines with AIML Operations team.
  4. Calibrate gate thresholds against production reality and maintain evaluation drift hygiene (golden-set rotation, hold-out sets, judge calibration).
  5. Produce the monthly quality report per agent: eval trends, failure-mode taxonomy, and defect clusters with reproduction traces.
  6. Mentor members of the team when needed and review their work for quality and consistency.

SECTION C: QUALIFICATIONS / EXPERIENCE / KNOWLEDGE REQUIRED

Indicate key knowledge and skills required for this role to perform the tasks to a satisfactory level. To also specify a suitable level of qualification required (i.e. basic, advanced, or professional), where applicable.

Category
Essential for this role
Good to have

Education and Qualifications
Bachelor’s or Master’s degree in Computer Science or a related field

Work Experience
· 6+ years in ML/data/software with strong evaluation or quality focus

• Hands-on experience evaluating LLM or ML systems
· Experience with agentic systems or RAG pipelines

Technical / Professional Skills

Please provideat least 3
· LLM/agent evaluation design and statistical rigour

· Python and evaluation tooling (promptfoo, DeepEval, or custom harnesses)

· Data analysis and metric interpretation

· Tracing/observability tooling
· Adversarial testing / red‑teaming

Non-Technical / Soft Skills
· Analytical rigour and attention to detail

· Clear technical writing for gate findings

 · Ability to influence build teams on quality
• Management presentation

Other Task-Specific Knowledge
· Understanding of telco customer intents and journeys
• Responsible-AI and safety evaluation practices

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Agent Evaluation Engineer
AI Agent Evaluation Engineer

Letitbex AI • Telangana

On-site
INR 4,000,000 - 7,500,000
AI Testing Specialist LLM Evaluation & Qualit
AI Testing Specialist LLM Evaluation & Qualit

Hucon Solutions • Hyderabad

On-site
INR 1,000,000 - 2,000,000
Applied AI Engineer
Applied AI Engineer

Infer • Karnataka

On-site
INR 1,500,000 - 2,500,000
LLM / Agentic Evaluation Rig Engineer
LLM / Agentic Evaluation Rig Engineer

PHIZENIX • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
Senior AI Evaluation & Reliability Engineer
Senior AI Evaluation & Reliability Engineer

Aubergine Solutions Pvt. Ltd. • Ahmedabad District

On-site
INR 3,000,000 - 6,000,000
Great Place To Work certified
QA Engineer
QA Engineer

E2M Solutions • Ahmedabad District

On-site
INR 700,000 - 1,200,000
27503 AI Agent Operations Engineer
27503 AI Agent Operations Engineer

Cephas Consultancy Services Private Limited • Pune District

On-site
INR 4,500,000 - 6,500,000
AI Evaluations Team Lead
AI Evaluations Team Lead

FNZ • Pune District

On-site
INR 9,541,000 - 12,405,000
Staff Software Development Test Engineer - AI Evaluation
Staff Software Development Test Engineer - AI Evaluation

Tekion Corp • Bengaluru

On-site
INR 900,000 - 1,300,000
Forward Deployed Engineer
Forward Deployed Engineer

Insight Global • Hyderabad

On-site
INR 3,500,000 - 6,500,000