AI QA Engineer – LLM, RAG & Prompt Testing

StatusNeo

Gurugram District

On-site

INR 2,500,000 - 3,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

StatusNeo is seeking an AI QA Engineer to rigorously test Generative AI, LLM-based applications, AI chatbots, and RAG systems. The role emphasizes AI response validation, hallucination detection, grounding verification, and prompt testing, going beyond manual testing or basic Selenium automation.

The ideal candidate will collaborate with product, engineering, and data science teams to define AI testing strategy, build evaluation suites, and ensure quality across multi-turn conversations and

Qualifications

  • 5+ years of QA / SDET / automation testing experience.
  • Hands-on experience in AI testing, GenAI testing, LLM testing, chatbot testing, or RAG testing.
  • Experience validating model outputs against expected outcomes, source documents, or evaluation metrics.
  • Familiarity with CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, or Azure DevOps).
  • Experience with Jira, Agile/Scrum, test planning, defect reporting, and regression testing.

Responsibilities

  • Test AI chatbots, LLM-powered applications, and GenAI workflows across functional, regression, API, and end-to-end scenarios.
  • Validate RAG retrieval results, including context relevance, retrieval accuracy, source grounding, and response completeness.
  • Evaluate AI-generated responses for hallucinations, factual accuracy, answer relevancy, consistency, and faithfulness.
  • Verify grounding and citations by checking whether model responses are supported by retrieved source documents.
  • Perform prompt validation, prompt robustness testing, jailbreak testing, and prompt injection attack testing.
  • Validate model responses against expected outcomes, golden datasets, business rules, and acceptance criteria.
  • Build or execute AI evaluation suites using tools/frameworks such as RAGAS, DeepEval, LangChain evaluation, LangSmith, prompt evaluation frameworks, or custom Python-based metrics.
  • Test multi-turn chatbot conversations for context retention, fallback handling, safety behaviour, and response consistency.
  • Collaborate with product, engineering, data science, and QA teams to define AI testing strategy and quality metrics.
  • Log AI defects clearly with prompt, context, source data, actual response, expected response, screenshots/logs, and evaluation reason.
  • Support automation using Python, Playwright, Selenium, REST API testing, Postman, GitHub Actions, Jenkins, or similar tools.

Skills

AI testing
GenAI testing
LLM testing
Chatbot testing
RAG testing
Automation testing

Education

Bachelor's degree in CS/IT or related

Tools

Playwright
Selenium
PyTest
REST Assured
Postman

Job description

We are looking for an AI QA Engineer with hands-on experience in testing Generative AI, LLM-based applications, AI chatbots, and RAG systems.

This role is not suitable for candidates with only manual testing or generic Selenium automation experience. The ideal candidate should have practical exposure to AI response validation, hallucination detection, prompt testing, grounding verification, RAG retrieval validation, and LLM evaluation frameworks.

Key Responsibilities

  • Test AI chatbots, LLM-powered applications, and GenAI workflows across functional, regression, API, and end-to-end scenarios.
  • Validate RAG retrieval results, including context relevance, retrieval accuracy, source grounding, and response completeness.
  • Evaluate AI-generated responses for hallucinations, factual accuracy, answer relevancy, consistency, and faithfulness.
  • Verify grounding and citations by checking whether model responses are supported by retrieved source documents.
  • Perform prompt validation, prompt robustness testing, jailbreak testing, and prompt injection attack testing.
  • Validate model responses against expected outcomes, golden datasets, business rules, and acceptance criteria.
  • Build or execute AI evaluation suites using tools/frameworks such as RAGAS, DeepEval, LangChain evaluation, LangSmith, prompt evaluation frameworks, or custom Python-based metrics.
  • Test multi-turn chatbot conversations for context retention, fallback handling, safety behaviour, and response consistency.
  • Collaborate with product, engineering, data science, and QA teams to define AI testing strategy and quality metrics.
  • Log AI defects clearly with prompt, context, source data, actual response, expected response, screenshots/logs, and evaluation reason.
  • Support automation using Python, Playwright, Selenium, REST API testing, Postman, GitHub Actions, Jenkins, or similar tools.

Must-Have Skills

  • 5+ years of QA / SDET / automation testing experience
  • Hands-on experience in AI testing, GenAI testing, LLM testing, chatbot testing, or RAG testing
  • Strong understanding of:
  • RAG retrieval validation
  • Experience validating model outputs against expected outcomes, source documents, golden datasets, or evaluation metrics
  • Experience with automation tools such as Playwright, Selenium, PyTest, or REST Assured
  • API testing experience using Postman / REST APIs
  • Familiarity with CI/CD pipelines such as GitHub Actions, Jenkins, GitLab CI, or Azure DevOps
  • Experience with Jira, Agile/Scrum, test planning, defect reporting, and regression testing

Good-to-Have Skills

  • Experience with RAGAS, DeepEval, LangSmith, Langfuse, LlamaIndex, LangChain, or LLM-as-a-Judge
  • Exposure to OpenAI, Azure OpenAI, Claude, Gemini, or other LLM APIs
  • Knowledge of vector databases such as FAISS, Pinecone, ChromaDB, or Weaviate
  • Understanding of embeddings, chunking, retrieval, context precision/recall, and faithfulness metrics
  • Experience testing agentic AI workflows, AI agents, MCP-based workflows, or multi-agent systems
  • Performance testing exposure for AI systems, including latency, TTFC, streaming consistency, and response quality.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI QA Engineer / Generative AI Test Engineer
AI QA Engineer / Generative AI Test Engineer

Cybage • Pune District

On-site
INR 900,000 - 1,300,000
Senior AI QA Engineer
Senior AI QA Engineer

Biz2X • Dadri

On-site
INR 1,500,000 - 2,100,000
Gen AI QA Engineer
Gen AI QA Engineer

Deqode • Bengaluru

On-site
INR 1,200,000 - 2,100,000
Senior AI Testing Engineer
Senior AI Testing Engineer

VidvanConnect Software Solutions Pvt. Ltd. • Delhi

On-site
INR 1,800,000 - 3,000,000
Senior Test Engineer - AI/ML Testing
Senior Test Engineer - AI/ML Testing

Information Technology • Bengaluru

On-site
INR 2,500,000 - 4,500,000
QA Engineer with AI
QA Engineer with AI

Infoya Inc. • Navalur

On-site
INR 900,000 - 1,500,000
Senior Test Analyst
Senior Test Analyst

Smartstream • Mumbai

On-site
INR 2,500,000 - 3,500,000
QA Engineer
QA Engineer

E2M Solutions • Ahmedabad District

On-site
INR 700,000 - 1,200,000
AI/LLM QA Engineer - Agentic and Multi-Agent System Testing
AI/LLM QA Engineer - Agentic and Multi-Agent System Testing

Crew Kraftorz LLP • Hyderabad

Hybrid
INR 1,500,000 - 2,200,000
Senior QA AI/ML Engineer
Senior QA AI/ML Engineer

Freecharge • Delhi, Gurugram District

Hybrid
INR 1,000,000 - 1,500,000