QA Automation - AI Testing

Statusneo Technology Consulting

Hyderabad

On-site

INR 1,500,000 - 2,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Statusneo Technology Consulting in Hyderabad seeks an AI QA Engineer with hands-on experience in testing Generative AI, LLM-based apps, AI chatbots, and RAG systems. The role requires practical exposure to AI response validation, hallucination detection, prompt testing, grounding verification, and evaluation frameworks.

The candidate will test AI chatbots and GenAI workflows across functional, regression, API, and end-to-end scenarios, validating RAG retrieval and model outputs, while

Qualifications

  • 5+ years QA/SDET/automation testing experience.
  • Hands-on AI testing, GenAI, LLM, chatbot, or RAG testing.
  • Hallucination detection and grounding validation concepts.
  • Prompt testing, jailbreak testing, and prompt injection familiarity.
  • RAG retrieval validation and AI model evaluation metrics.
  • Experience validating outputs against sources or golden datasets.
  • Programming in Python, JavaScript, or TypeScript.
  • Automation tools including Playwright, Selenium, PyTest, REST Assured.
  • API testing with Postman / REST APIs; CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, Azure DevOps).
  • Jira, Agile/Scrum, test planning, defect reporting, and regression testing.

Responsibilities

  • Test AI chatbots, LLM-powered apps, and GenAI workflows across functional, regression, API, and end-to-end tests.
  • Validate RAG retrieval results: relevance, grounding, source accuracy, and response completeness.
  • Evaluate AI-generated responses for hallucinations, factual accuracy, relevance, and faithfulness.
  • Verify grounding and citations by checking alignment with retrieved sources.
  • Conduct prompt validation, robustness, jailbreak testing, and prompt injection testing.
  • Assess model responses against expected outcomes, golden datasets, and business rules.
  • Build or run AI evaluation suites using frameworks like RAGAS, DeepEval, LangChain evaluation, LangSmith, or similar.
  • Test multi-turn conversations for context retention, fallbacks, safety, and consistency.
  • Collaborate with product, engineering, data science, and QA to define AI testing strategy and quality metrics.
  • Log AI defects with prompt, context, source data, actual vs. expected responses, and evaluation rationale.
  • Support automation using Python, Playwright, Selenium, REST API testing, Postman, and CI/CD tooling.

Skills

QA / SDET experience
AI testing
LLM testing
Bot testing
RAG testing
Python
JavaScript/TypeScript
Playwright
Selenium
Postman / REST APIs
CI/CD (GitHub Actions/Jenkins)
Jira/Agile/Scrum

Tools

Playwright
Selenium
PyTest
REST Assured
GitHub Actions
Jenkins
GitLab CI
Azure DevOps

Job description

Role & responsibilities
About the Role

We are looking for an AI QA Engineer with hands‑on experience in testing Generative AI, LLM-based applications, AI chatbots, and RAG systems.

This role is not suitable for candidates with only manual testing or generic Selenium automation experience. The ideal candidate should have practical exposure to AI response validation, hallucination detection, prompt testing, grounding verification, RAG retrieval validation, and LLM evaluation frameworks.


Key Responsibilities
  • Test AI chatbots, LLM-powered applications, and GenAI workflows across functional, regression, API, and end‑to‑end scenarios.
  • Validate RAG retrieval results, including context relevance, retrieval accuracy, source grounding, and response completeness.
  • Evaluate AI‑generated responses for hallucinations, factual accuracy, answer relevancy, consistency, and faithfulness.
  • Verify grounding and citations by checking whether model responses are supported by retrieved source documents.
  • Perform prompt validation, prompt robustness testing, jailbreak testing, and prompt injection attack testing.
  • Validate model responses against expected outcomes, golden datasets, business rules, and acceptance criteria.
  • Build or execute AI evaluation suites using tools/frameworks such as RAGAS, DeepEval, LangChain evaluation, LangSmith, prompt evaluation frameworks, or custom Python‑based metrics.
  • Test multi‑turn chatbot conversations for context retention, fallback handling, safety behaviour, and response consistency.
  • Collaborate with product, engineering, data science, and QA teams to define AI testing strategy and quality metrics.
  • Log AI defects clearly with prompt, context, source data, actual response, expected response, screenshots/logs, and evaluation reason.
  • Support automation using Python, Playwright, Selenium, REST API testing, Postman, GitHub Actions, Jenkins, or similar tools.

Must‑Have Skills
  • 5+ years of QA / SDET / automation testing experience
  • Hands‑on experience in AI testing, GenAI testing, LLM testing, chatbot testing, or RAG testing
  • Strong understanding of:
    • Hallucination detection
    • Grounding validation
    • Prompt testing
    • Prompt injection / jailbreak testing
    • RAG retrieval validation
    • AI model response evaluation
  • Experience validating model outputs against expected outcomes, source documents, golden datasets, or evaluation metrics
  • Working knowledge of Python, JavaScript, or TypeScript
  • Experience with automation tools such as Playwright, Selenium, PyTest, or REST Assured
  • API testing experience using Postman / REST APIs
  • Familiarity with CI/CD pipelines such as GitHub Actions, Jenkins, GitLab CI, or Azure DevOps
  • Experience with Jira, Agile/Scrum, test planning, defect reporting, and regression testing

Good‑to‑Have Skills
  • Experience with RAGAS, DeepEval, LangSmith, Langfuse, LlamaIndex, LangChain, or LLM‑as‑a‑Judge
  • Exposure to OpenAI, Azure OpenAI, Claude, Gemini, or other LLM APIs
  • Knowledge of vector databases such as FAISS, Pinecone, ChromaDB, or Weaviate
  • Understanding of embeddings, chunking, retrieval, context precision/recall, and faithfulness metrics
  • Experience testing agentic AI workflows, AI agents, MCP‑based workflows, or multi‑agent systems
  • Performance testing exposure for AI systems, including latency, TTFC, streaming consistency, and response quality
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI QA Engineer – LLM, RAG & Prompt Testing
AI QA Engineer – LLM, RAG & Prompt Testing

StatusNeo • Gurugram District

On-site
INR 2,500,000 - 3,500,000
Senior AI QA Engineer
Senior AI QA Engineer

Biz2X • Dadri

On-site
INR 1,500,000 - 2,100,000
Gen AI QA Engineer
Gen AI QA Engineer

Deqode • Bengaluru

On-site
INR 1,200,000 - 2,100,000
Automation Testing-AI
Automation Testing-AI

Sonata Software • Hyderabad, Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 2,500,000
QA Engineer
QA Engineer

YO IT Consulting • Mumbai

Hybrid
INR 800,000 - 1,400,000
AI Test Automation Specialist
AI Test Automation Specialist

PwC • Bengaluru

On-site
INR 800,000 - 1,600,000
Quality Assurance Lead
Quality Assurance Lead

Cloud Angles Digital Transformation • Hyderabad

On-site
INR 2,600,000 - 3,800,000
Quality Assurance Engineer
Quality Assurance Engineer

Valiance Solutions • Dadri

On-site
INR 800,000 - 1,500,000
Senior QA AI/ML Engineer
Senior QA AI/ML Engineer

Freecharge • Delhi, Gurugram District

Hybrid
INR 1,000,000 - 1,500,000
QA Engineer
QA Engineer

E2M Solutions • Ahmedabad District

On-site
INR 700,000 - 1,200,000