AI/ML Quality Engineer

Allegis Global Solutions (AGS)

Bengaluru

On-site

INR 1,800,000 - 3,800,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Allegis Global Solutions in Bengaluru, India is seeking an AI Quality Engineer to design and implement testing for AI agents, conversational interfaces, and agentic workflows. You will define non-deterministic output testing, build evaluation frameworks, and measure groundedness, factuality, and task completion.

You will integrate automated tests into CI/CD, develop test harnesses for LLM systems, and monitor production quality with dashboards.

Qualifications

  • 3-7 years of software engineering or quality engineering experience.
  • Strong programming skills in Python and/or TypeScript—you write test code, not just test cases.
  • Experience designing and building automated test frameworks.
  • Understanding of AI/ML systems—you know why testing LLM outputs is different from testing deterministic code.
  • Experience with CI/CD pipelines and integrating automated tests into build processes.
  • Ability to reason about non-deterministic systems and design meaningful quality metrics.
  • Strong analytical skills—you can look at agent outputs and determine whether they're good enough.

Responsibilities

  • Define testing strategies for AI agents, conversational interfaces, and agentic workflows
  • Design behavioral test suites for non-deterministic outputs - where "correct" isn't binary
  • Build evaluation frameworks that measure groundedness, factuality, relevance, and task completion
  • Identify failure modes specific to AI systems: hallucinations, prompt injection, context window limitations, drift
  • Develop testing approaches for each architecture pattern: RAG, function calling, human-in-the-loop, autonomous workflows
  • Build automated evaluation pipelines that run as part of CI/CD
  • Create test harnesses for LLM-based systems - mocking, fixtures, and reproducible test scenarios
  • Develop regression suites that detect quality degradation when prompts, models, or data change
  • Build monitoring and alerting for production agent quality (accuracy, latency, error rates)
  • Maintain test infrastructure: test data management, environment setup, reporting dashboards

Skills

Python
TypeScript
Automated test frameworks
CI/CD
Non-deterministic systems reasoning
Analytical skills
AI/ML systems understanding

Tools

pytest
DeepEval
Ragas
Docker
Git
CI/CD pipelines

Job description

Responsibilities
Testing Strategy & Design
  • Define testing strategies for AI agents, conversational interfaces, and agentic workflows
  • Design behavioral test suites for non-deterministic outputs - where "correct" isn't binary
  • Build evaluation frameworks that measure groundedness, factuality, relevance, and task completion
  • Identify failure modes specific to AI systems: hallucinations, prompt injection, context window limitations, drift
  • Develop testing approaches for each architecture pattern: RAG, function calling, human-in-the-loop, autonomous workflows
Test Automation & Infrastructure
  • Build automated evaluation pipelines that run as part of CI/CD
  • Create test harnesses for LLM-based systems - mocking, fixtures, and reproducible test scenarios
  • Develop regression suites that detect quality degradation when prompts, models, or data change
  • Build monitoring and alerting for production agent quality (accuracy, latency, error rates)
  • Maintain test infrastructure: test data management, environment setup, reporting dashboards
Evaluation & Metrics
  • Define quality metrics for each solution - what to measure and what thresholds matter
  • Build and maintain evaluation datasets (ground truth, reference outputs, edge case collections)
  • Conduct systematic prompt evaluation when prompts or models change
  • Track quality trends over time and identify when re-evaluation is needed
  • Report quality metrics to the team and stakeholders in clear, actionable terms
Collaboration & Quality Culture
  • Partner with AI Solutions Engineers to define testability requirements during design
  • Work with AI Solutions Analysts to translate acceptance criteria into test scenarios
  • Review solution designs from a quality and testability perspective
  • Advocate for quality practices across the team - testing isn't an afterthought, it's part of delivery
  • Contribute to incident response by diagnosing quality failures and building regression tests
Qualifications
Required
  • 3-7 years of software engineering or quality engineering experience
  • Strong programming skills in Python and/or TypeScript—you write test code, not just test cases
  • Experience designing and building automated test frameworks
  • Understanding of AI/ML systems—you know why testing LLM outputs is different from testing deterministic code
  • Experience with CI/CD pipelines and integrating automated tests into build processes
  • Ability to reason about non-deterministic systems and design meaningful quality metrics
  • Strong analytical skills—you can look at agent outputs and determine whether they're good enough
Preferred
  • Experience testing AI/ML applications, conversational interfaces, or chatbots
  • Background in LLM evaluation: prompt testing, groundedness scoring, factuality checking
  • Familiarity with evaluation frameworks (DeepEval, Ragas, custom evaluation pipelines)
  • Experience with Microsoft Power Platform (Power Automate, Copilot Studio) testing
  • Background in Azure services and cloud-based test infrastructure
  • Experience with load testing and performance testing for API-based systems
  • Familiarity with staffing, HR tech, or workforce management domains
Technology Stack
  • Languages: Python, TypeScript
  • Platforms: Azure (Container Apps, Functions, AI Services), Microsoft 365
  • Testing: pytest, evaluation frameworks (DeepEval, Ragas, custom), load testing tools
  • AI/ML: LLM evaluation, prompt testing, RAG evaluation, behavioral testing
  • Data: REST APIs, Dataverse, SQL
  • Tools: Git, GitHub, CI/CD pipelines, Docker, monitoring/alerting (Application Insights)

We don't expect expertise in everything. AI quality engineering is a new discipline - we expect strong engineering fundamentals and the ability to figure out new problems.

What We're NOT Looking For
  • Manual testers who write test cases in spreadsheets
  • QA professionals who treat testing as a gate at the end of development rather than a practice woven into it
  • People who expect deterministic pass/fail for every test - AI quality requires nuance
  • Engineers who test to the spec but don't think about how real users will break things
What Makes You Stand Out
  • You've tested a system where "correct" was hard to define - and found a way to measure it anyway
  • You write test code that's as clean and maintainable as production code
  • You think about edge cases that nobody else considers
  • You can explain why a particular quality metric matters and what threshold makes sense
  • You've built test automation that actually caught regressions before they hit production
  • You're comfortable saying "this isn't good enough" and backing it up with data
What We're Building

The AI Engineering team delivers intelligent solutions for AGS's global clients:

  • Intelligent Agents - Conversational AI that helps hiring managers, recruiters, and internal teams get work done faster
  • Agentic Workflows - Automated processes where AI executes tasks with human oversight
  • Foundation Capabilities - Reusable AI services that power multiple solutions

You'll make sure these systems work reliably - not just at launch, but as models change, data evolves, and usage scales.

Career Growth

AI quality engineering is an emerging discipline with no ceiling. Growth paths include:

  • Depth - Become the team's authority on AI evaluation and testing methodology, influencing quality standards across the organization
  • Breadth - Move into a Senior or Lead AI Solutions Engineer role, bringing your quality mindset to architecture and delivery
  • Specialization - Build expertise in areas like LLM security testing, AI safety, or evaluation research
Why Join Us
  • Solve new problems - AI testing is an unsolved discipline. You'll invent approaches, not just follow playbooks
  • Broad visibility - Touch every solution the team builds and understand the full architecture
  • Write real code - Build evaluation frameworks and test infrastructure, not just test cases
  • Production systems - Test agents serving Fortune 500 clients, where quality actually matters
  • Shape quality culture - Define what quality means for AI at AGS and build the systems to enforce it
  • Growing field - AI quality engineering is early. Getting good at it now puts you ahead of the industry
About Allegis Global Solutions

We are founded on a culture that is passionate about transforming the way the world acquires talent by delivering client-focused solutions that make a difference for businesses worldwide. From refining how businesses manage their contingent workforce to strengthening employer brands to recruit top talent, our integrated solutions drive business results. As an industry leader, we draw upon decades of experience to design innovative tools, products, and processes. We develop competitive practices that position organizations for growth and we deliver the insight needed to succeed in today's global marketplace. As a workplace, we focus on relationships - with each other, our clients, and our candidates. In fact, serving others is one of our core values. We support open communication and recognize that giving constructive criticism can be even harder than receiving it. We appreciate the fearless and the passionate, who force us to be better. Everything we do sits on a pillar of diversity - diverse perspectives, backgrounds, and ideas drive innovation and make us successful. See what it's like to work at AGS by searching #LifeAtAGS on any social network.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Quality Engineer
AI Quality Engineer

Allegis Global Solutions • Bengaluru

On-site
INR 1,200,000 - 1,800,000
AI Solutions Engineer 1
AI Solutions Engineer 1

Allegis Global Solutions • Bengaluru

On-site
INR 600,000 - 1,000,000
AI Quality Engineer
AI Quality Engineer

Allegis Group Services, Inc. • India

On-site
INR 1,200,000 - 2,400,000
AI Solutions Engineer
AI Solutions Engineer

Allegis Group Services, Inc. • India

On-site
INR 900,000 - 1,800,000
AI Solutions Engineer
AI Solutions Engineer

Allegis Global Solutions • Bengaluru

On-site
INR 1,200,000 - 2,400,000
AI Solutions Analyst
AI Solutions Analyst

Allegis Group Services, Inc. • India

On-site
INR 600,000 - 900,000
AI Solutions Analyst
AI Solutions Analyst

Allegis Global Solutions • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Quality Engineer - GenAI and Data Platforms
Quality Engineer - GenAI and Data Platforms

Optum India • Hyderabad

On-site
INR 700,000 - 1,100,000
Quality Engineer - GenAI and Data Platforms
Quality Engineer - GenAI and Data Platforms

Optum • Hyderabad

On-site
INR 600,000 - 1,200,000
Senior AI QA Engineer
Senior AI QA Engineer

Biz2X • Dadri

On-site
INR 1,500,000 - 2,100,000