GenAI AI/ML Test Engineer: Evaluation & Benchmarking

Improving

Hinoba-an

On-site

PHP 796,000 - 1,194,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this recruiter — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Improving is seeking an AI/ML Test Engineer – GenAI to design and automate evaluation strategies for Agentic AI applications in Mumbai. You will develop datasets, test cases, and benchmark suites to measure agent performance, reasoning quality, tool usage, and workflow effectiveness.

The role emphasizes identifying hallucinations, biases, safety risks, and failure patterns while collaborating with AI Engineers and product teams to improve reliability and specify actionable recommendations.

Qualifications

  • Strong experience in Generative AI, LLMs, and Agentic AI systems.
  • Hands-on expertise with AI evaluation frameworks (RAGAS, DeepEval, TruLens, LangSmith, Promptfoo, etc.)
  • Proficiency in Python and AI/ML development libraries.
  • Knowledge of Prompt Engineering, prompt testing, and optimization.
  • Ability to define and track evaluation metrics such as accuracy, relevance, groundedness, hallucination rate, latency, and user satisfaction.
  • Experience in creating automated evaluation pipelines and benchmarking frameworks.
  • Strong understanding of AI safety, guardrails, bias testing, and responsible AI practices.
  • Familiarity with REST APIs, JSON, vector databases, and knowledge retrieval systems.
  • Experience in A/B testing, human-in-the-loop evaluation, and red teaming.
  • Strong experience in Manual Testing of AI/GenAI applications, including functional, exploratory, UAT, regression, and end-to-end testing.
  • Expertise in validating Agent Reasoning, Tool Calling, Workflow Execution, and Response Quality.
  • Hands-on experience in Automation Testing using Python frameworks.

Responsibilities

  • Design, execute, and automate evaluation strategies for Agentic AI applications.
  • Develop evaluation datasets, test cases, and benchmark suites.
  • Measure and improve agent performance, reasoning quality, tool usage, and workflow effectiveness.
  • Analyze model outputs and identify hallucinations, biases, safety risks, and failure patterns.
  • Collaborate with AI Engineers, Product Teams, and Domain Experts to improve agent quality and reliability.
  • Generate evaluation reports, dashboards, and actionable recommendations.

Skills

Generative AI
LLMs
Agentic AI systems
AI evaluation frameworks
Python
Prompt engineering
Evaluation metrics
Automation testing
Agent reasoning
Tool calling
Workflow execution
REST APIs
JSON
Vector databases
A/B testing
Human-in-the-loop evaluation
Red teaming
Quality assessment

Tools

RAGAS
DeepEval
TruLens
LangSmith
Promptfoo

Job description

Improving is seeking an AI/ML Test Engineer – GenAI to design and automate evaluation strategies for Agentic AI applications in Mumbai. You will develop datasets, test cases, and benchmark suites to measure agent performance, reasoning quality, tool usage, and workflow effectiveness.

The role emphasizes identifying hallucinations, biases, safety risks, and failure patterns while collaborating with AI Engineers and product teams to improve reliability and specify actionable recommendations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Test Engineer – GenAI
AI/ML Test Engineer – GenAI

Improving • Hinoba-an

On-site
PHP 796,000 - 1,194,000
AI Quality Assurance Engineer - Automation & GenAI
AI Quality Assurance Engineer - Automation & GenAI

blaseek • Manila

On-site
PHP 3,616,000 - 4,823,000
AI Automation Tester (Hybrid Setup)
AI Automation Tester (Hybrid Setup)

blaseek • Manila

On-site
PHP 3,616,000 - 4,823,000
AI Benchmarking & Performance Engineer
AI Benchmarking & Performance Engineer

DevRev • Metro Manila

On-site
PHP 1,200,000 - 2,400,000
QA Automation Engineer with AI
QA Automation Engineer with AI

Infoya Inc. • Hinoba-an

On-site
PHP 600,000 - 1,000,000
Lead AI Evaluation & Safety Architect
Lead AI Evaluation & Safety Architect

BrightClaim • Philippines

On-site
PHP 1,800,000 - 2,400,000
AI-Enabled Automation Engineer for GenAI & Full-Stack
AI-Enabled Automation Engineer for GenAI & Full-Stack

BlackCube Labs • Hinoba-an

Hybrid
PHP 789,000 - 1,577,000
GenAI QA Engineer – AI/ML Testing & Automation
GenAI QA Engineer – AI/ML Testing & Automation

Accenture • Hinoba-an

On-site
PHP 900,000 - 1,500,000
Lead AI Engineer: Evaluation & Safety Architect
Lead AI Engineer: Evaluation & Safety Architect

Latitude • Philippines

On-site
PHP 2,000,000 - 3,400,000
AI Architect
AI Architect

Career Connect • Philippines

Hybrid
PHP 1,800,000 - 3,000,000