AI Quality Engineer

Allegis Group Services, Inc.

India

On-site

INR 1,200,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Allegis Group Services, Inc. is seeking an AI Quality Engineer to design testing strategies for intelligent agents and non-deterministic outputs. You will build evaluation pipelines, automated testing frameworks, and CI/CD integrations to ensure production reliability across AI solutions.

The role spans across multiple architectures, including RAG and function calling, with emphasis on groundedness, factuality, and latency metrics, requiring strong Python/TypeScript coding skills.

Qualifications

  • 3–7 years of software engineering or quality engineering experience
  • Strong programming skills in Python and/or TypeScript—you write test code, not just test cases
  • Experience designing and building automated test frameworks
  • Understanding of AI/ML systems—you know why testing LLM outputs is different from testing deterministic code
  • Experience with CI/CD pipelines and integrating automated tests into build processes
  • Ability to reason about non-deterministic systems and design meaningful quality metrics
  • Strong analytical skills—you can look at agent outputs and determine whether they're good enough

Responsibilities

  • Define testing strategies for AI agents, conversational interfaces, and agentic workflows
  • Design behavioral test suites for non-deterministic outputs—where "correct" isn't binary
  • Build evaluation frameworks that measure groundedness, factuality, relevance, and task completion
  • Identify failure modes specific to AI systems: hallucinations, prompt injection, context window limitations, drift
  • Develop testing approaches for each architecture pattern: RAG, function calling, human-in-the-loop, autonomous workflows
  • Build automated evaluation pipelines that run as part of CI/CD
  • Create test harnesses for LLM-based systems—mocking, fixtures, and reproducible test scenarios
  • Develop regression suites that detect quality degradation when prompts, models, or data change
  • Build monitoring and alerting for production agent quality (accuracy, latency, error rates)
  • Define quality metrics for each solution—what to measure and thresholds matter
  • Build and maintain evaluation datasets (ground truth, reference outputs, edge case collections)
  • Conduct systematic prompt evaluation when prompts or models change
  • Track quality trends over time and identify when re-evaluation is needed
  • Report quality metrics to the team and stakeholders in clear, actionable terms
  • Partner with AI Solutions Engineers to define testability requirements during design
  • Work with AI Solutions Analysts to translate acceptance criteria into test scenarios
  • Review solution designs from a quality and testability perspective
  • Advocate for quality practices across the team—testing isn't an afterthought, it's part of delivery
  • Contribute to incident response by diagnosing quality failures and building regression tests

Skills

Python
TypeScript
Test automation
CI/CD pipelines
LLM testing
Quality metrics

Tools

pytest
Power Platform
Azure
CI/CD tools

Job description

Job Description

About the Role

Testing AI systems is a fundamentally different problem than testing traditional software. Outputs are non-deterministic. "Correct" is often a spectrum. And the failure modes—hallucinations, drift, prompt injection—don't show up in unit tests. We need an engineer who understands this and can build the testing strategies, evaluation frameworks, and quality infrastructure to keep our agents reliable in production.

As an AI Quality Engineer, you'll design how we test intelligent agents, agentic workflows, and Foundation Layer capabilities. This is not a manual QA role—you'll write code, build evaluation pipelines, and create automated testing frameworks that run in CI/CD. You'll define what "quality" means for AI systems at AGS and build the systems to measure it.

You'll work across every solution the team builds, which means you'll have broad visibility into the architecture and deep understanding of how our agents behave in the real world. If you're an engineer who cares about quality and wants to solve testing problems that most teams haven't figured out yet, this is the role.

Responsibilities

  • Define testing strategies for AI agents, conversational interfaces, and agentic workflows
  • Design behavioral test suites for non-deterministic outputs—where "correct" isn't binary
  • Build evaluation frameworks that measure groundedness, factuality, relevance, and task completion
  • Identify failure modes specific to AI systems: hallucinations, prompt injection, context window limitations, drift
  • Develop testing approaches for each architecture pattern: RAG, function calling, human-in-the-loop, autonomous workflows

Test Automation & Infrastructure

  • Build automated evaluation pipelines that run as part of CI/CD
  • Create test harnesses for LLM-based systems—mocking, fixtures, and reproducible test scenarios
  • Develop regression suites that detect quality degradation when prompts, models, or data change
  • Build monitoring and alerting for production agent quality (accuracy, latency, error rates)

Evaluation & Metrics

  • Define quality metrics for each solution—what to measure and what thresholds matter
  • Build and maintain evaluation datasets (ground truth, reference outputs, edge case collections)
  • Conduct systematic prompt evaluation when prompts or models change
  • Track quality trends over time and identify when re-evaluation is needed
  • Report quality metrics to the team and stakeholders in clear, actionable terms

Collaboration & Quality Culture

  • Partner with AI Solutions Engineers to define testability requirements during design
  • Work with AI Solutions Analysts to translate acceptance criteria into test scenarios
  • Review solution designs from a quality and testability perspective
  • Advocate for quality practices across the team—testing isn't an afterthought, it's part of delivery
  • Contribute to incident response by diagnosing quality failures and building regression tests
Qualifications

Qualifications

Required

  • 3–7 years of software engineering or quality engineering experience
  • Strong programming skills in Python and/or TypeScript—you write test code, not just test cases
  • Experience designing and building automated test frameworks
  • Understanding of AI/ML systems—you know why testing LLM outputs is different from testing deterministic code
  • Experience with CI/CD pipelines and integrating automated tests into build processes
  • Ability to reason about non-deterministic systems and design meaningful quality metrics
  • Strong analytical skills—you can look at agent outputs and determine whether they're good enough

Preferred

  • Experience testing AI/ML applications, conversational interfaces, or chatbots
  • Background in LLM evaluation: prompt testing, groundedness scoring, factuality checking
  • Familiarity with evaluation frameworks (DeepEval, Ragas, custom evaluation pipelines)
  • Experience with Microsoft Power Platform (Power Automate, Copilot Studio) testing
  • Background in Azure services and cloud-based test infrastructure
  • Experience with load testing and performance testing for API-based systems
  • Familiarity with staffing, HR tech, or workforce management domains
  • Languages: Python, TypeScript
  • Testing: pytest, evaluation frameworks (DeepEval, Ragas, custom), load testing tools
  • Data: REST APIs, Dataverse, SQL

We don't expect expertise in everything. AI quality engineering is a new discipline—we expect strong engineering fundamentals and the ability to figure out new problems.

What We're NOT Looking For

  • Manual testers who write test cases in spreadsheets
  • QA professionals who treat testing as a gate at the end of development rather than a practice woven into it
  • People who expect deterministic pass/fail for every test—AI quality requires nuance
  • Engineers who test to the spec but don't think about how real users will break things

What Makes You Stand Out

  • You've tested a system where "correct" was hard to define—and found a way to measure it anyway
  • You write test code that's as clean and maintainable as production code
  • You think about edge cases that nobody else considers
  • You can explain why a particular quality metric matters and what threshold makes sense
  • You've built test automation that actually caught regressions before they hit production
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Test Engineer - AI/ML Testing
Senior Test Engineer - AI/ML Testing

Information Technology • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Quality Assurance Engineer
Quality Assurance Engineer

Valiance Solutions • Dadri

On-site
INR 800,000 - 1,500,000
QA Engineer
QA Engineer

E2M Solutions • Ahmedabad District

On-site
INR 700,000 - 1,200,000
Senior Quality Automation Engineer — AI-Native (APAC)
Senior Quality Automation Engineer — AI-Native (APAC)

Right Stakes • India

On-site
INR 1,200,000 - 1,800,000
Senior AI QA Engineer
Senior AI QA Engineer

Biz2X • Dadri

On-site
INR 1,500,000 - 2,100,000
Senior Test Analyst
Senior Test Analyst

Smartstream • Mumbai

On-site
INR 2,500,000 - 3,500,000
AI QA Engineer
AI QA Engineer

Deqode • Ahmedabad District

On-site
INR 600,000 - 800,000
Senior AI Testing Engineer
Senior AI Testing Engineer

VidvanConnect Software Solutions Pvt. Ltd. • Delhi

On-site
INR 1,800,000 - 3,000,000
Quality Engineering consultant
Quality Engineering consultant

Artech L.L.C. • Bengaluru

On-site
INR 600,000 - 1,200,000
QA Engineer
QA Engineer

YO IT Consulting • Mumbai

Hybrid
INR 800,000 - 1,400,000