Agent Evaluation & Instrumentation Engineer

NCS Group

Pune District

On-site

INR 1,800,000 - 3,000,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NCS Group in Pune, India seeks an AI Evaluation & Instrumentation Engineer to design, calibrate, and adjudicate evaluation suites for agent archetypes. You will lead gate reviews, own regression-pack methodology, and ensure quality and drift hygiene across AI/ML models.

Key tasks include offline evaluation suites, operability reviews, model-update regression, threshold calibration, monthly quality reports, and mentoring the team. Strong background in ML evaluation and telco AI is required.

Qualifications

  • Bachelor's or master's degree in CS or related field.
  • 6+ years in ML/data/software with strong evaluation or quality focus.
  • Hands-on experience evaluating LLM or ML systems.
  • Experience with agentic systems or RAG pipelines.
  • Telco AI domain.

Responsibilities

  • Design and maintain offline evaluation suites (golden sets, regression packs, adversarial/safety probes) and the continuous-evaluation scoring pipeline.
  • Conduct operability gate reviews: assess evidence packs, reproduce evaluation results, and recommend go/no-go with documented findings.
  • Own the model-update regression pack methodology and adjudicate regression runs against archived baselines with AIML Operations team.
  • Calibrate gate thresholds against production reality and maintain evaluation drift hygiene.
  • Produce the monthly quality report per agent: eval trends, failure-mode taxonomy, and defect clusters with reproduction traces.
  • Mentor members of the team when needed and review their work for quality and consistency.

Skills

LLM/agent evaluation design
Python and evaluation tooling
Data analysis and metric interpretat
Tracing/observability tooling
Adversarial testing / red-teaming

Education

Bachelor's or Master's degree in Computer Science or related field

Tools

promptfoo
DeepEval
custom harnesses

Job description

Job Description - AI Evaluation & Instrumentation Engineer
SECTION A: POSITION SUMMARY

State the objective and purpose of the role.

Serve as the technical evaluator within the team, responsible for the design, calibration, and adjudication of evaluation suites and gate thresholds across agent archetypes. Lead gate reviews under the Team Lead s authority, own the regression-pack methodology for AI/ML model changes, and act as the technical custodian of evaluation quality and drift hygiene.

SECTION B: KEY RESPONSIBILITIES AND RESULTS

Indicate key responsibilities and performance indicators of this role.

For existing role, please indicate additional responsibilities in bold.
  • 1. Design and maintain offline evaluation suites (golden sets, regression packs, adversarial/safety probes) and the continuous-evaluation scoring pipeline across archetypes.
  • 2. Conduct operability gate reviews: assess evidence packs, reproduce evaluation results, and recommend go/no-go with documented findings.
  • 3. Own the model-update regression pack methodology and adjudicate regression runs against archived baselines with AIML Operations team.
  • 4. Calibrate gate thresholds against production reality and maintain evaluation drift hygiene (golden-set rotation, hold-out sets, judge calibration).
  • 5. Produce the monthly quality report per agent: eval trends, failure-mode taxonomy, and defect clusters with reproduction traces.
  • 6. Mentor members of the team when needed and review their work for quality and consistency.
SECTION C: QUALIFICATIONS / EXPERIENCE / KNOWLEDGE REQUIRED

Indicate key knowledge and skills required for this role to perform the tasks to a satisfactory level. To also specify a suitable level of qualification required (i.e. basic, advanced, or professional), where applicable.

Education and Qualifications
  • Bachelor s or Master s degree in Computer Science or a related field
Work Experience
  • 6+ years in ML/data/software with strong evaluation or quality focus
  • - Hands-on experience evaluating LLM or ML systems
  • Experience with agentic systems or RAG pipelines
  • Telco AI domain
Technical / Professional Skills

Please provide at least 3

  • LLM/agent evaluation design and statistical rigour
  • Python and evaluation tooling (promptfoo, DeepEval, or custom harnesses)
  • Data analysis and metric interpretation
  • Tracing/observability tooling
  • Adversarial testing / red-teaming
Non-Technical / Soft Skills
  • Analytical rigour and attention to detail
  • Clear technical writing for gate findings
  • Ability to influence build teams on quality
  • Management presentation
Other Task-Specific Knowledge
  • Understanding of telco customer intents and journeys - Responsible-AI and safety evaluation practices

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

67503 Agent Evaluation & Instrumentation Engineer
67503 Agent Evaluation & Instrumentation Engineer

Cephas Consultancy Services Private Limited • Pune District

On-site
INR 3,000,000 - 4,500,000
AI Agent Evaluation Engineer
AI Agent Evaluation Engineer

Letitbex AI • Telangana

On-site
INR 4,000,000 - 7,500,000
AI Quality Assurance Engineer
AI Quality Assurance Engineer

Thirdeye Data Inc. • Bengaluru

Hybrid
INR 1,200,000 - 2,000,000
Senior AI Evaluation & Reliability Engineer
Senior AI Evaluation & Reliability Engineer

Aubergine Solutions Pvt. Ltd. • Ahmedabad District

On-site
INR 3,000,000 - 6,000,000
Great Place To Work certified
Senior GenAI Engineer (AI Evaluation Engineer)
Senior GenAI Engineer (AI Evaluation Engineer)

FM India • Bengaluru

On-site
INR 1,800,000 - 3,000,000
AI Evaluation Engineer - Proofline
AI Evaluation Engineer - Proofline

Fermi AI • Bengaluru

On-site
INR 1,500,000 - 2,100,000
AI Engineer
AI Engineer

CodeRound AI • Delhi

On-site
INR 1,800,000 - 3,200,000
Staff Software Development Test Engineer - AI Evaluation
Staff Software Development Test Engineer - AI Evaluation

Tekion Corp • Bengaluru

On-site
INR 900,000 - 1,300,000
Lead AI Engineer - Agentic Engineering
Lead AI Engineer - Agentic Engineering

Blend360 India • Hyderabad

Hybrid
INR 3,000,000 - 5,000,000
Regulatory Compliance GxP AI Model Consultant
Regulatory Compliance GxP AI Model Consultant

EY • India

Hybrid
INR 1,500,000 - 2,500,000