AI/ML Test Engineer – GenAI

Improving

Hinoba-an

On-site

PHP 796,000 - 1,194,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Improving is seeking an AI/ML Test Engineer – GenAI to design and automate evaluation strategies for Agentic AI applications in Mumbai. You will develop datasets, test cases, and benchmark suites to measure agent performance, reasoning quality, tool usage, and workflow effectiveness.

The role emphasizes identifying hallucinations, biases, safety risks, and failure patterns while collaborating with AI Engineers and product teams to improve reliability and specify actionable recommendations.

Qualifications

  • Strong experience in Generative AI, LLMs, and Agentic AI systems.
  • Hands-on expertise with AI evaluation frameworks (RAGAS, DeepEval, TruLens, LangSmith, Promptfoo, etc.)
  • Proficiency in Python and AI/ML development libraries.
  • Knowledge of Prompt Engineering, prompt testing, and optimization.
  • Ability to define and track evaluation metrics such as accuracy, relevance, groundedness, hallucination rate, latency, and user satisfaction.
  • Experience in creating automated evaluation pipelines and benchmarking frameworks.
  • Strong understanding of AI safety, guardrails, bias testing, and responsible AI practices.
  • Familiarity with REST APIs, JSON, vector databases, and knowledge retrieval systems.
  • Experience in A/B testing, human-in-the-loop evaluation, and red teaming.
  • Strong experience in Manual Testing of AI/GenAI applications, including functional, exploratory, UAT, regression, and end-to-end testing.
  • Expertise in validating Agent Reasoning, Tool Calling, Workflow Execution, and Response Quality.
  • Hands-on experience in Automation Testing using Python frameworks.

Responsibilities

  • Design, execute, and automate evaluation strategies for Agentic AI applications.
  • Develop evaluation datasets, test cases, and benchmark suites.
  • Measure and improve agent performance, reasoning quality, tool usage, and workflow effectiveness.
  • Analyze model outputs and identify hallucinations, biases, safety risks, and failure patterns.
  • Collaborate with AI Engineers, Product Teams, and Domain Experts to improve agent quality and reliability.
  • Generate evaluation reports, dashboards, and actionable recommendations.

Skills

Generative AI
LLMs
Agentic AI systems
AI evaluation frameworks
Python
Prompt engineering
Evaluation metrics
Automation testing
Agent reasoning
Tool calling
Workflow execution
REST APIs
JSON
Vector databases
A/B testing
Human-in-the-loop evaluation
Red teaming
Quality assessment

Tools

RAGAS
DeepEval
TruLens
LangSmith
Promptfoo

Job description

Improving is committed to building a great place to work by cultivating an environment that fosters professional and personal relationships. We value open communication, personal growth, and shared rewards, which result in sustainable success.

Voted “best place to work” numerous times, Improving strives to create and maintain a culture that exemplifies teamwork, excellence, and fun! We believe this kind of culture encourages both the inspiration and the motivation to achieve amazing things.

Title: AI/ML Test Engineer – GenAI

Location: Mumbai

Experience - 1-3 years

Technical Skills
  • Strong experience in Generative AI, LLMs, and Agentic AI systems
  • Hands-on expertise with AI evaluation frameworks (RAGAS, DeepEval, TruLens, LangSmith, Promptfoo, etc.)
  • Proficiency in Python and AI/ML development libraries
  • Knowledge of Prompt Engineering, prompt testing, and optimization
  • Ability to define and track evaluation metrics such as accuracy, relevance, groundedness, hallucination rate, latency, and user satisfaction
  • Experience in creating automated evaluation pipelines and benchmarking frameworks
  • Strong understanding of AI safety, guardrails, bias testing, and responsible AI practices
  • Familiarity with REST APIs, JSON, vector databases, and knowledge retrieval systems
  • Experience in A/B testing, human-in-the-loop evaluation, and red teaming
  • Strong experience in Manual Testing of AI/GenAI applications, including functional, exploratory, UAT, regression, and end-to-end testing
  • Expertise in validating Agent Reasoning, Tool Calling, Workflow Execution, and Response Quality
  • Hands-on experience in Automation Testing using Python frameworks
Key Responsibilities
  • Design, execute, and automate evaluation strategies for Agentic AI applications.
  • Develop evaluation datasets, test cases, and benchmark suites.
  • Measure and improve agent performance, reasoning quality, tool usage, and workflow effectiveness.
  • Analyze model outputs and identify hallucinations, biases, safety risks, and failure patterns.
  • Collaborate with AI Engineers, Product Teams, and Domain Experts to improve agent quality and reliability.
  • Generate evaluation reports, dashboards, and actionable recommendations.
About Improving

Improving is a modern digital services company dedicated to positively changing the perception of the IT professional. We offer innovative solutions through consulting, software development, and training to help thousands of our clients achieve new heights in a competitive and ever-changing market.

As our company continues to grow, we are looking for enthusiastic thought leaders to join our team. Improving has a unique mix of passionate professionals who strive to grow and thrive in new ways. We are committed to establishing and maintaining an inclusive culture that allows all Improvers to bring their authentic selves to work each day. This is why we work hard to build inclusion and diversity in our workplace, so we can all do amazing things and succeed together.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI AI/ML Test Engineer: Evaluation & Benchmarking
GenAI AI/ML Test Engineer: Evaluation & Benchmarking

Improving • Hinoba-an

On-site
PHP 796,000 - 1,194,000
QA Automation Engineer with AI
QA Automation Engineer with AI

Infoya Inc. • Hinoba-an

On-site
PHP 600,000 - 1,000,000
Lead AI Engineer 4C
Lead AI Engineer 4C

Latitude • Philippines

On-site
PHP 2,000,000 - 3,400,000
Lead AI Engineer 4C
Lead AI Engineer 4C

BrightClaim • Philippines

On-site
PHP 1,800,000 - 2,400,000
AI Architect
AI Architect

Career Connect • Philippines

Hybrid
PHP 1,800,000 - 3,000,000
Lead Architect
Lead Architect

V2 Solutions • Hinoba-an

Hybrid
PHP 1,194,000 - 2,122,000
AI Engineer
AI Engineer

Dry Ground • Philippines

On-site
PHP 2,995,000 - 4,794,000
Competitive salary and performance-based incentives
Flexible work environment
Collaborative innovation-driven culture
AI / ML Engineer
AI / ML Engineer

Accenture • Hinoba-an

On-site
PHP 1,322,000 - 1,984,000
Director Associate Distinguished Engineer (Solution Architect - GenAI + RAG + Agentic AI)
Director Associate Distinguished Engineer (Solution Architect - GenAI + RAG + Agentic AI)

Nagarro • Hinoba-an

On-site
PHP 2,400,000 - 4,200,000
GenAI Engineer - Hybrid Role with Growth & Benefits
GenAI Engineer - Hybrid Role with Growth & Benefits

Thinking Machines Data Science • Philippines

Hybrid
PHP 800,000 - 1,200,000
Hybrid Set-Up
Health benefits
Professional development budget
+1