Agentic QA

Horizon Industries International Limited

Dadri

On-site

INR 4,000,000 - 7,000,000

Full time

1 hour ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Agentic AI Testing, Evaluation & Automation is seeking an experienced QA leader in India to define and drive testing strategies for GenAI and autonomous AI systems. You will validate agent behavior, evaluation metrics, and end-to-end workflows, building scalable evaluation pipelines and dashboards.

The role requires deep experience in AI testing methodologies, hands-on automation with Python and Playwright, and collaboration with Product, Engineering, Data Science, and AI Research teams to

Qualifications

  • 7–12 years of experience in Software Testing, Quality Engineering, or Test Automation.
  • Minimum 3+ years of hands-on experience in GenAI, LLM Testing, Agentic AI Testing, or AI Quality Engineering.
  • Strong understanding of LLMs, AI agents, RAG, prompt validation, tool calling, agent memory, MCP, and multi-agent orchestration.
  • Experience defining AI quality metrics, evaluation methodologies, benchmarking frameworks, and model comparison approaches.
  • Hands-on automation experience with Python, Playwright, Pytest, API automation, test framework development, and CI/CD quality gates.
  • Exposure to cloud platforms such as Azure, AWS, or GCP.

Responsibilities

  • Define and execute testing strategies for LLM-based, multi-agent, RAG, and Agentic AI systems.
  • Validate autonomous agent behavior, reasoning, memory, tool usage, API/database integrations, and end-to-end workflows.
  • Evaluate AI outputs for accuracy, relevance, groundedness, consistency, completeness, toxicity, bias, hallucination risk, and guardrail compliance.
  • Define AI quality KPIs such as hallucination rate, groundedness score, agent success rate, task completion rate, response relevancy, latency, cost efficiency, and user satisfaction.
  • Build automated evaluation pipelines, quality scoring mechanisms, dashboards, and CI/CD-integrated quality gates.
  • Develop reusable test harnesses, simulators, and benchmarking frameworks to compare models, prompts, and agent configurations.

Skills

GenAI Testing
LLM Testing
AI Quality Engineering
Test Automation
Python
Playwright
Pytest
CI/CD Gates
AI Evaluation Metrics

Tools

DeepEval
Ragas
LangSmith
OpenAI Evals
LangGraph
CrewAI
AutoGen
Semantic Kernel

Job description

Agentic AI Testing, Evaluation & Automation
  • Define and execute testing strategies for LLM-based, multi-agent, RAG, and Agentic AI systems.
  • Validate autonomous agent behavior, reasoning, memory, tool usage, API/database integrations, and end-to-end workflows.
  • Evaluate AI outputs for accuracy, relevance, groundedness, consistency, completeness, toxicity, bias, hallucination risk, and guardrail compliance.
  • Define AI quality KPIs such as hallucination rate, groundedness score, agent success rate, task completion rate, response relevancy, latency, cost efficiency, and user satisfaction.
  • Build automated evaluation pipelines, quality scoring mechanisms, dashboards, and CI/CD-integrated quality gates.
  • Develop reusable test harnesses, simulators, and benchmarking frameworks to compare models, prompts, and agent configurations.
Team Leadership & Capability Building
  • Build and lead a team of Agentic AI Quality Engineers.
  • Define team structure, testing standards, best practices, and governance models.
  • Mentor QA engineers in AI testing methodologies, evaluation techniques, and automation frameworks.
  • Drive innovation and adoption of emerging AI testing tools and technologies.
  • Collaborate with Product, Engineering, Data Science, and AI Research teams to improve overall AI quality.
Reporting & Stakeholder Management
  • Provide quality assessments and recommendations to leadership and stakeholders.
  • Present testing outcomes, risk assessments, KPI trends, and model evaluation reports.
  • Drive quality governance for Agentic AI initiatives across the organization.
  • Ensure traceability of testing activities, evaluation criteria, and quality benchmarks.
Required Skills & Experience
Technical Skills
  • 7–12 years of experience in Software Testing, Quality Engineering, or Test Automation.
  • Minimum 3+ years of hands‑on experience in GenAI, LLM Testing, Agentic AI Testing, or AI Quality Engineering.
  • Strong understanding of LLMs, AI agents, RAG, prompt validation, tool calling, agent memory, MCP, and multi‑agent orchestration.
  • Experience defining AI quality metrics, evaluation methodologies, benchmarking frameworks, and model comparison approaches.
  • Hands‑on automation experience with Python, Playwright, Pytest, API automation, test framework development, and CI/CD quality gates.
  • Experience with AI evaluation frameworks such as DeepEval, Ragas, LangSmith, OpenAI Evals, or equivalent tools.
  • Exposure to cloud platforms such as Azure, AWS, or GCP.
Soft Skills
  • Strong communication and stakeholder management skills.
  • Analytical mindset with strong problem‑solving ability.
  • Self‑driven, outcome‑oriented, and capable of leading multiple initiatives in a fast‑evolving AI ecosystem.
Preferred Qualifications
  • Experience testing enterprise Agentic AI platforms and autonomous AI systems.
  • Hands‑on exposure to frameworks or tools such as LangGraph, CrewAI, AutoGen, Semantic Kernel, Microsoft Copilot Studio, TruLens, or Promptfoo.
  • Exposure to AI observability, monitoring, model governance, responsible AI, and AI safety practices.
  • Experience building AI quality dashboards and KPI reporting systems.
  • ISTQB, AI Testing, GenAI, Azure AI, AWS AI, or equivalent certifications.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead QA - Agentic AI
Lead QA - Agentic AI

Sycamore Informatics Inc. • India

On-site
INR 3,500,000 - 6,000,000
AL/ML Engineer
AL/ML Engineer

Qentelli • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Qentelli Solutions - AI Quality Assurance Engineer
Qentelli Solutions - AI Quality Assurance Engineer

Qentelli • Hyderabad

On-site
INR 2,200,000 - 3,600,000
AI Testing Specialist LLM Evaluation & Qualit
AI Testing Specialist LLM Evaluation & Qualit

Hucon Solutions • Hyderabad

On-site
INR 1,000,000 - 2,000,000
AI Quality Engineer
AI Quality Engineer

Allegis Group Services, Inc. • India

On-site
INR 1,200,000 - 2,400,000
AI/LLM QA Engineer - Agentic and Multi-Agent System Testing
AI/LLM QA Engineer - Agentic and Multi-Agent System Testing

Crew Kraftorz LLP • Hyderabad

Hybrid
INR 1,500,000 - 2,200,000
AI Test Engineer
AI Test Engineer

Qentelli • Hyderabad

On-site
INR 1,800,000 - 2,400,000
Gen AI QA Engineer
Gen AI QA Engineer

Deqode • Bengaluru

On-site
INR 1,200,000 - 2,100,000
AI Engineer
AI Engineer

Qentelli • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Quality Assurance Lead
Quality Assurance Lead

Cloud Angles Digital Transformation • Hyderabad

On-site
INR 2,600,000 - 3,800,000