Senior AI QA engineer

Fulcrum Worldwide Software

Pune District

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fulcrum Digital is seeking a Senior AI QA Engineer in Pune to design and execute QA for LLM agents, RAG pipelines, and AI-assisted decision tools. You will validate outputs against ground truth, identify hallucinations, and perform cross-model evaluations across GPT, Claude, Gemini, and Perplexity.

Strong Jira reporting and regression testing discipline are required. The role emphasizes evaluating AI reasoning, building evaluation rubrics for non-deterministic outputs, and ensuring quality in

Qualifications

  • 7–8 years of QA experience with focus on Generative AI/LLM-based projects.
  • Hands-on testing of chatbots, RAG systems, or agentic AI pipelines.
  • Ability to perform ground truth validation and detect hallucinations in model outputs.
  • Familiarity with multi-model evaluation, prompt-aware testing, and Jira-based defect reporting.

Responsibilities

  • Design and execute test cases for LLM agents, RAG pipelines, agentic workflows.
  • Validate AI outputs against ground truth using structured accuracy scoring.
  • Detect hallucinations, reasoning gaps, and source fabrication in model content.
  • Run multi-model comparative testing across GPT, Claude, Gemini, and Perplexity for accuracy and latency.
  • Test prompt versions iteratively and track accuracy changes across cycles.
  • Validate citation accuracy and cross-document context handling.
  • Design edge cases and negative tests for AI-specific failure modes.
  • Perform regression testing after model upgrades and maintain QA sign-off in Jira.

Skills

QA experience
Gen AI/LLM testing
Ground truth validation
JIRA defect reporting
Multi-model evaluation

Tools

JIRA
Azure OpenAI
AWS Bedrock
SharePoint AI

Job description

Urgent - Immediate Hiring - Senior AI QA Engineer - Fulcrum Digital, Pune (Work mode - 2 days WFO)

Fulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing.

Key Responsibilities
  • Design and execute test cases for LLM agents, RAG pipelines, agentic workflows, and AI-assisted decision tools
  • Validate AI outputs against ground truth using structured accuracy scoring (NAICS, risk exposure flags, hazard group mapping)
  • Detect hallucinations, reasoning gaps, source fabrication, and misattributions in model-generated content
  • Run multi-model comparative testing across GPT, Claude, Gemini, and Perplexity evaluating accuracy, latency, and output completeness
  • Test prompt versions iteratively and track accuracy changes across prompt cycles
  • Validate citation accuracy, document ingestion pipelines, and cross-document context handling
  • Design edge case and negative tests for AI-specific failure modes content filter triggers, tool call limits, missing documents, and incomplete synthesis
  • Perform regression testing after model upgrades, prompt changes, and backend fixes, and maintain structured QA sign-off in JIRA
What Makes This Role Different from Traditional QA
  • You evaluate whether an AI is reasoning correctly — not just whether the UI behaves as expected
  • You build evaluation rubrics for non-deterministic outputs and apply LLM-as-a-Judge techniques to score quality at scale
  • You treat every model or prompt change as a potential accuracy regression, not just a functional one
  • You understand that in live AI systems, a passing test today does not guarantee a passing test tomorrow
Required Skillsets
  • 7–8 years of QA experience with minimum 2 years in Generative AI / LLM-based projects
  • Hands-on experience testing chatbots, RAG systems, or agentic AI pipelines
  • Proven ability to perform ground truth validation and detect hallucinations and reasoning failures
  • Familiarity with multi-model evaluation, prompt-aware testing, and JIRA-based defect reporting
Preferred Skillsets
  • Background in insurance or regulated industries; exposure to underwriting or risk classification concepts
  • Familiarity with Azure OpenAI, AWS Bedrock, or SharePoint-integrated AI environments
  • Knowledge of AI governance, content filtering, and PII redaction validation
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior QA AI/ML Engineer
Senior QA AI/ML Engineer

Freecharge • Delhi, Gurugram District

Hybrid
INR 1,000,000 - 1,500,000
AI Testing Engineer
AI Testing Engineer

Black Turtle • Bengaluru, Pune District

Hybrid
INR 1,500,000 - 3,000,000
Senior AI QA Engineer
Senior AI QA Engineer

Biz2X • Dadri

On-site
INR 1,500,000 - 2,100,000
Senior Test Analyst
Senior Test Analyst

Smartstream • Mumbai

On-site
INR 2,500,000 - 3,500,000
Principal QA – AI & Conversational Systems (Pune)
Principal QA – AI & Conversational Systems (Pune)

Codvo Private Limited • Pune District

On-site
INR 1,500,000 - 2,000,000
AI QA Engineer – LLM, RAG & Prompt Testing
AI QA Engineer – LLM, RAG & Prompt Testing

StatusNeo • Gurugram District

On-site
INR 2,500,000 - 3,500,000
Generative AI Engineer
Generative AI Engineer

Tekskills Inc. • Pune District

On-site
INR 1,200,000 - 1,600,000
QA Engineer
QA Engineer

YO IT Consulting • Mumbai

Hybrid
INR 800,000 - 1,400,000
QA Analyst with AI
QA Analyst with AI

Tekskills • Bulandshahr

Hybrid
INR 800,000 - 1,100,000
QA Analyst with AI
QA Analyst with AI

Tekskills • Bengaluru

Hybrid
INR 1,100,000 - 1,500,000