Automation Testing-AI

Sonata Software

Hyderabad, Chennai District, Bengaluru

Hybrid

INR 1,200,000 - 2,500,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Sonata Software in Hyderabad seeks an AI QA / LLM Automation Engineer to design, develop, and execute testing strategies for generative AI systems, including LLMs, RAG pipelines, and autonomous agents.

You will build Python automation, integrate tests into CI/CD pipelines, define non-deterministic evaluation metrics, and collaborate with AI researchers and product managers to improve accuracy, context retention, and reliability.

Qualifications

  • Proficient in Python with clean, scalable automation scripts.
  • Deep understanding of LLM architectures and prompt engineering.
  • Experience testing RAG pipelines and agentic frameworks.
  • Experience validating non-deterministic outputs with semantic evaluation.
  • Familiar with CI/CD tools and version control.

Responsibilities

  • AI Model Validation: Design and execute test strategies for evaluating LLM responses with focus on accuracy and reducing hallucinations.
  • AI Agent Testing: Test the logic, state management, and tool-calling capabilities of autonomous AI agents.
  • Python Automation: Develop custom automated testing frameworks using Python to support CI/CD pipelines.
  • CI/CD Integration: Integrate testing frameworks into existing CI/CD pipelines for ongoing evaluation.
  • Evaluation Metrics: Define metrics for non-deterministic AI features where exact-match is insufficient.
  • Cross-Functional Collaboration: Work with AI developers, data scientists, and product managers to set quality benchmarks.

Skills

Python
LLM architectures
RAG pipelines
Agent frameworks
CI/CD basics
Non-deterministic testing
Problem solving

Tools

LangChain
LlamaIndex
Git
GitHub Actions
GitLab CI
Jenkins

Job description

Role & responsibilities
Job Description: AI QA / LLM Automation Engineer
Job Overview

We are seeking a highly skilled and innovative AI QA / LLM Automation Engineer to join our team. In this pivotal role, you will be responsible for designing, developing, and executing sophisticated testing strategies tailored specifically for generative AI systems. You will play a critical role in evaluating Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) pipelines, and autonomous AI agents to ensure high accuracy, contextual relevance, and robust logical reasoning.

Key Roles & Responsibilities
  • AI Model Validation: Design and execute comprehensive test strategies specifically for evaluating LLM responses. Focus heavily on measuring accuracy, ensuring long-term context retention, and drastically reducing or mitigating hallucinations.
  • AI Agent Testing: Rigorously test the underlying logic, complex state management, and tool-calling capabilities of autonomous AI agents across various edge cases.
  • Python Automation: Develop, maintain, and scale custom automated testing frameworks from scratch using Python.
  • CI/CD Integration: Continuously evaluate AI models by seamlessly integrating custom testing and evaluation frameworks into existing CI/CD pipelines.
  • Evaluation Metrics: Define and implement novel evaluation metrics for non-deterministic AI features where standard deterministic software testing (exact-match) falls short.
  • Cross-Functional Collaboration: Work closely with AI developers, data scientists, and product managers to establish quality benchmarks for AI-driven features.
Required Skills & Qualifications
  • Programming: Strong programming proficiency in Python, with an emphasis on writing clean, scalable, and maintainable automation scripts.
  • AI/ML Knowledge: Deep understanding of Large Language Model (LLM) architectures, prompt engineering, and model behaviors.
  • Framework Experience: Hands-on experience building or testing RAG (Retrieval-Augmented Generation) pipelines and agentic frameworks (e.g., LangChain, LlamaIndex).
  • Non-Deterministic Testing: Proven experience in validating non-deterministic outputs where traditional exact-match assertions do not apply (e.g., semantic similarity, LLM-as-a-judge, or custom evaluation heuristics).
  • Automation & DevOps: Solid understanding of CI/CD tools (e.g., GitHub Actions, GitLab CI, Jenkins) and version control (Git).
  • Problem Solving: Strong analytical mindset with the ability to foresee edge cases in generative AI applications.
  • Key Responsibilities:
  • AI Model Validation: Design and execute test strategies specifically for evaluating LLM responses, focusing on accuracy, context retention, and reducing hallucinations.
  • AI Agent Testing: Test the logic, state management, and tool-calling capabilities of autonomous AI agents.
  • Python Automation: Develop custom automated testing frameworks using Python to continuously evaluate AI models in CI/CD pipelines.
  • Required Skills:
  • Strong programming proficiency in Python.
  • Deep understanding of LLM architectures, RAG (Retrieval-Augmented Generation), and agentic frameworks (e.g., LangChain, LlamaIndex).
  • Experience validating non-deterministic outputs where traditional exact-match assertions do not apply.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GenAI QA / AI Test Engineer
GenAI QA / AI Test Engineer

Halcer • Bengaluru

On-site
INR 1,800,000 - 2,700,000
Senior QA AI Testing Engineer
Senior QA AI Testing Engineer

NewVision • Hyderabad

On-site
INR 2,500,000 - 5,000,000
AI Test Engineer
AI Test Engineer

Deqode • Bengaluru

On-site
INR 900,000 - 1,500,000
Senior Quality Assurance AI Engineer II CloudAngles
Senior Quality Assurance AI Engineer II CloudAngles

Cloud Angles Digital Transformation • Hyderabad

Hybrid
INR 1,200,000 - 1,800,000
AI/LLM QA Engineer - Agentic and Multi-Agent System Testing
AI/LLM QA Engineer - Agentic and Multi-Agent System Testing

Crew Kraftorz LLP • Hyderabad

Hybrid
INR 1,500,000 - 2,200,000
Principal Software Engineer
Principal Software Engineer

Cadence Design Systems • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Lead Javascript Automation Test Engineer with AI & Agentic Testing
Lead Javascript Automation Test Engineer with AI & Agentic Testing

EPAM Systems • Maharashtra

On-site
INR 4,000,000 - 7,000,000
Lead Javascript Automation Test Engineer with AI & Agentic Testing
Lead Javascript Automation Test Engineer with AI & Agentic Testing

EPAM Systems • Chennai District

On-site
INR 2,500,000 - 4,200,000
Lead Javascript Automation Test Engineer with AI & Agentic Testing
Lead Javascript Automation Test Engineer with AI & Agentic Testing

EPAM Systems • Coimbatore District

On-site
INR 2,500,000 - 5,000,000
Hiring AI QA Automation Lead | Agentic AI | LLMs | Playwright
Hiring AI QA Automation Lead | Agentic AI | LLMs | Playwright

Vibehackers • Hyderabad

On-site
INR 2,500,000 - 3,800,000