AI Testing Specialist LLM Evaluation & Qualit

Hucon Solutions

Hyderabad

On-site

INR 1,000,000 - 2,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hucon Solutions is seeking a Testing Specialist focused on LLM evaluation and quality assurance in Hyderabad. You will design evaluation frameworks with RAGAS and custom scorers, build regression tests, and automate quality gates within CI/CD pipelines.

You will collaborate with AI/ML engineers and DevSecOps teams in Germany and Hyderabad to ensure compliance with AI governance and industry standards. Strong Python, pytest, and scripting skills are required, as is experience with LLM tooling and

Qualifications

  • Bachelors in CS/Engineering or related field with AI/ML testing experience.
  • Experience testing AI/ML systems or LLM-based applications.
  • Strong Python and pytest skills with test automation experience.

Responsibilities

  • Design and implement model evaluation frameworks using RAGAS, custom scorers, and automated accuracy benchmarking.
  • Develop adversarial prompt testing suites to probe agent behaviour for hallucinations and edge cases.
  • Build regression tests to verify agent behaviour across updates and prompts.
  • Integrate AI test tooling into CI/CD pipelines with GitHub Actions.
  • Collaborate with AI/ML Engineers and DevSecOps teams to raise quality for production releases.

Skills

Python
pytest
test automation
scripting
LLM evaluation
RAGAS
adversarial testing
regression testing
CI/CD
GitHub Actions
LangGraph
LangChain
retrieval augmented generation
vector DB validation
Docker
Kubernetes

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

RAGAS
GitHub Actions
LangGraph
LangChain
Docker
Kubernetes

Job description

Testing Specialist - LLM Evaluation & Quality

Design and implement model evaluation frameworks using RAGAS, custom

scorers, and automated accuracy benchmarking to validate agent outputs

against structured specifications and Gherkin acceptance criteria.

Develop adversarial prompt testing suites that probe agent behaviour for

hallucinations, prompt injection vulnerabilities, and edge-case failures across the

Refinement, Decision, and Coding Agents.

Build and maintain regression test suites that continuously verify agent behaviour

consistency across model updates, prompt changes, and skill-file modifications

within the LangGraph-based architecture.

Integrate AI-specific test tooling into CI/CD pipelines using GitHub Actions,

ensuring every agent deployment is gated by automated quality, fairness, and

reliability checks aligned with EU AI Act and DORA requirements.

Define and track quantitative quality metrics, confidence thresholds, retrieval

precision, response latency, and drift detection, providing actionable dashboards

to the engineering team.

Collaborate closely with AI/ML Engineers and DevSecOps teams in Germany

and Hyderabad to establish testing standards, review evaluation results, and

continuously raise the quality bar for production agent releases.

Bachelor's degree in Computer Science, Engineering, or a related field, with

demonstrated experience in testing AI/ML systems or LLM-based applications.

Strong programming skills in Python with hands‑on proficiency in pytest, test

automation frameworks, and scripting for evaluation pipelines.

Practical experience with LLM evaluation methodologies and tools such as

RAGAS, custom scoring functions, accuracy benchmarking, and adversarial/red

team testing techniques.

Solid understanding of CI/CD integration using GitHub Actions or comparable

platforms, with the ability to embed AI test gates into automated deployment

workflows.

Familiarity with agentic architectures such as LangGraph or LangChain, including

knowledge of retrieval‑augmented generation (RAG) patterns and multi‑agent

orchestration concepts. Nice‑to‑have: Experience with regulatory compliance testing (EU AI Act, DORA),

familiarity with vector database validation and containerization

(Docker/Kubernetes).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/LLM QA Engineer - Agentic and Multi-Agent System Testing
AI/LLM QA Engineer - Agentic and Multi-Agent System Testing

Crew Kraftorz LLP • Hyderabad

Hybrid
INR 1,500,000 - 2,200,000
QA Engineer
QA Engineer

YO IT Consulting • Mumbai

Hybrid
INR 800,000 - 1,400,000
Automation Testing-AI
Automation Testing-AI

Sonata Software • Hyderabad, Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 2,500,000
Senior QA AI Testing Engineer
Senior QA AI Testing Engineer

NewVision • Hyderabad

On-site
INR 2,500,000 - 5,000,000
Senior AI Testing Engineer (Generative AI)
Senior AI Testing Engineer (Generative AI)

Wfnen • Bengaluru

Hybrid
INR 1,200,000 - 1,600,000
GenAI / Agent Testing Engineer(Bangalore only)
GenAI / Agent Testing Engineer(Bangalore only)

PwC • Bengaluru

On-site
INR 1,200,000 - 2,400,000
QA Engineer
QA Engineer

E2M Solutions • Ahmedabad District

On-site
INR 700,000 - 1,200,000
AI Data & Quality Analytics Tester(GenAI / Agent Testing Engineer)
AI Data & Quality Analytics Tester(GenAI / Agent Testing Engineer)

PwC • Bengaluru

On-site
INR 1,200,000 - 1,800,000
AI Quality Engineer | Contractual | Hyderabad
AI Quality Engineer | Contractual | Hyderabad

Side • Hyderabad

On-site
INR 502,000 - 837,000
AL/ML Engineer
AL/ML Engineer

Qentelli • Hyderabad

On-site
INR 4,000,000 - 7,000,000