Hiring || Gen AI Quality Assurance (BLR/Pune)

2coms

Bengaluru

On-site

INR 2,500,000 - 4,000,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

2coms in Bengaluru is seeking an experienced Quality Engineer to lead the validation of Generative AI applications, including chatbots and RAG-based systems. This role focuses on ensuring reliability and accuracy through automated testing frameworks and API evaluation.

The ideal candidate will have 7–12 years in software QA with AI/ML, strong Python scripting, and familiarity with ARIZE, CI/CD, and evaluation methodologies. This position offers opportunities to influence AI quality at scale.

Qualifications

  • 7–12 years in software QA with AI/ML focus.
  • Experience designing automated test frameworks using Python and pytest.
  • Hands-on using evaluation tools such as DeepEval and RAGAS.

Responsibilities

  • Execute testing for AI agents including chatbots, RAG-driven systems, summarizers, and note analyzers.
  • Validate AI outputs against accuracy, relevance, consistency and bias.
  • Architect and maintain automated evaluation pipelines with DeepEval and RAGAS.
  • Develop LLM-as-a-Judge protocols with clear rubrics and scoring.
  • Engineer pytest scripts to test API interactions and prompt responses.
  • Integrate test suites with CI/CD for local and production environments.
  • Capture metrics like hallucinations, faithfulness, and contextual accuracy.

Skills

Python
pytest
AI QA
Test automation
LLM-evaluation
LLM-as-a-Judge
CI/CD

Tools

DeepEval
RAGAS
Arize platform

Job description

Hiring || Gen AI Quality Assurance (BLR/Pune)
  • Dial in Extension to connect with Recruiter 410
Job Description
Gen AI Quality Assurance

Summary

We are seeking an experienced Quality Engineer with 7 to 12 years of background to lead the validation of advanced Generative AI applications. This role focuses on ensuring the reliability and accuracy of diverse AI agents, including chatbots, Retrieval-Augmented Generation (RAG) systems, and various summarization tools. The successful candidate will architect and deploy robust automated testing frameworks utilizing tools like pytest, DeepEval, and RAGAS to assess agent performance. Proficiency in the Arize platform is highly desirable. The position involves rigorous evaluation of AI interactions via APIs, employing both standard automation protocols and innovative LLM-as-a-Judge methodologies to measure critical quality indicators such as hallucination rates, faithfulness, contextual precision, and bias.

Responsibilities

  • Execute comprehensive testing strategies for a wide range of AI agents, including conversational chatbots, RAG-driven systems, document summarizers, and call note analyzers.
  • Validate AI-generated outputs against key performance standards, ensuring accuracy, relevance, consistency, and completeness while identifying issues like hallucinations or bias.
  • Architect and maintain automated evaluation pipelines leveraging the DeepEval and RAGAS frameworks to create reusable test suites for RAG performance, summarization fidelity, and groundedness.
  • Develop and refine LLM-as-a-Judge evaluation protocols, establishing clear rubrics and custom scoring mechanisms to objectively assess AI responses.
  • Engineer scalable automation scripts using pytest to validate API interactions, test prompt-response dynamics, and manage regression testing for AI models.
  • Integrate automated test suites with CI/CD pipelines, ensuring seamless execution in both local and production environments.
  • Capture and analyze detailed quality metrics, including toxicity levels, answer correctness, and overall response reliability, to drive continuous improvement.
  • Parameterize test cases to accommodate diverse prompts, contexts, and expected outcomes, facilitating batch evaluations of large test datasets.
Requirements
  • 7 to 12 years of professional experience in software quality assurance, with a specialized focus on AI and Machine Learning systems.
  • Proven expertise in designing, developing, and executing automated test frameworks, specifically using Python and pytest.
  • Hands-on experience implementing evaluation frameworks such as DeepEval and RAGAS for assessing Generative AI quality.
  • Strong understanding of LLM-as-a-Judge techniques and the ability to define custom evaluation criteria and rubrics.
  • Experience testing AI agents that communicate via APIs, including validation of request and response structures.
  • Familiarity with the Arize platform is considered a significant advantage.
  • Ability to analyze and interpret complex metrics related to faithfulness, contextual recall, hallucination, and bias.
  • Experience integrating test automation into CI/CD workflows and managing batch processing for test datasets.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Hiring | Gen AI Quality Assurance (BLR/Pune)
Hiring | Gen AI Quality Assurance (BLR/Pune)

2COMS Consulting Pvt. Ltd. • Bengaluru

On-site
INR 450,000 - 750,000
Senior AI Testing Engineer (Generative AI)
Senior AI Testing Engineer (Generative AI)

Wfnen • Bengaluru

On-site
INR 1,200,000 - 1,600,000
AI Data & Quality Analytics Tester
AI Data & Quality Analytics Tester

PwC • Bengaluru

On-site
INR 2,600,000 - 4,800,000
Functional AI Tester - GenAI
Functional AI Tester - GenAI

Michelin • Pune District

On-site
INR 1,200,000 - 1,800,000
Generative AI Engineer
Generative AI Engineer

Tekskills Inc. • Pune District

On-site
INR 1,200,000 - 1,600,000
Senior Quality Assurance AI Engineer II CloudAngles
Senior Quality Assurance AI Engineer II CloudAngles

Cloud Angles Digital Transformation • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Agentic AI Test Engineer
Agentic AI Test Engineer

Deutsche Telekom Digital Labs • Gurugram District

On-site
INR 1,620,000 - 1,980,000
AI Evaluation & Test Engineer
AI Evaluation & Test Engineer

BharatGen • Mumbai

On-site
INR 1,500,000 - 2,500,000
QA Engineer with AI
QA Engineer with AI

Infoya Inc. • Navalur

On-site
INR 900,000 - 1,500,000
QA Engineer
QA Engineer

YO IT Consulting • Mumbai

On-site
INR 800,000 - 1,400,000