AI Quality Assurance Engineer

Thirdeye Data Inc.

Bengaluru

Hybrid

INR 1,200,000 - 2,000,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Thirdeye Data Inc. in Bengaluru, India, is seeking an AI Quality Assurance Engineer to own the evaluation of our AI agent stack—models, prompts, tools, memory, and orchestration logic.

You will test all components, build tests, and measure with numbers to prevent regressions. You will establish evaluation infrastructure, KPIs, and regression datasets, ensuring tests cover multi-modality (text, image, audio, structured documents) and are version-controlled.

Qualifications

  • 4+ years of hands-on work in evaluation of LLM systems and agents.

Responsibilities

  • Own the evaluation process for the full AI agent stack and ensure readiness for deployment.
  • Build and maintain evaluation suites covering prompts, tools, memory, and orchestration logic.
  • Define quality metrics such as task success rate, latency, cost, and tool-call accuracy.
  • Build automated evaluation pipelines and integrate them with CI/CD.
  • Calibrate LLM judges against human labels and maintain tracing/logs for agent runs.

Skills

LLM evaluation
AI/Agent testing
Python
QA automation
Metrics definition
Multi-modality testing

Tools

pytest
Playwright
Selenium
GitHub Actions
Jenkins
GitLab CI
Postman
REST-assured
OpenTelemetry

Job description

AI Quality Assurance Engineer


Agent Evaluation

Location


Bengaluru, India (Onsite / Hybrid)

Department


AI Engineering / Quality

Reports to


Head of AI Engineering / Director of QA

Employment type


Full-time

About the Role

You will own the evaluation process for our AI agent system. The system includes models, prompts, tools, memory, and orchestration logic. You will test all of these parts.

You will build the tests that decide if an agent is ready to ship. You will find regressions before customers find them. You will measure agent quality with numbers, not opinions.

What You Will Do

Own the evaluation process

  • Build and maintain evaluation suites for the full agent stack.
  • Test prompt behavior, tool selection, tool calls, multi-step plans, memory use, error recovery, and task results.
  • Define quality metrics for each capability. Examples: task success rate, trajectory correctness, tool-call accuracy, latency, and cost.
  • Build and maintain golden datasets, regression suites, and adversarial test sets. Put these datasets under version control.

Build evaluation infrastructure

  • Build automated evaluation pipelines. Connect them to CI/CD.
  • Make sure each change to a model, prompt, tool, or the harness must pass evaluation before release.
  • Use more than one grading method: rule-based checks, exact match, semantic match, human review, and LLM-as-judge.
  • Calibrate LLM judges against human labels. Monitor the judges for drift.
  • Add tracing and logs to agent runs. Record each step, each tool call, and each cost. Make failures easy to diagnose.

Establish and standardize KPIs and metrics

  • Define the KPI set for agent quality. Include accuracy, faithfulness, answer relevance, context precision, context recall, tool-call correctness, hallucination rate, latency, and cost.
  • Write standard definitions for each metric. Write standard rubrics and report formats.
  • Make sure teams can compare results across models, prompts, and releases.
  • Extend evaluation to more than one modality: text, image, audio, and structured documents. Select the correct grading method for each modality.

Analyze failures and drive fixes

  • Do structured error analysis. Group failures by type.
  • Find the source of each failure: the model, the harness, or a tool.
  • Set sample sizes, pass thresholds, and confidence bounds. Agent runs are not deterministic. Plan for variance.
  • Work with engineers to reproduce failures and to verify fixes.

Protect the evaluation results

  • Find and prevent test contamination, overfitting to benchmarks, and metric gaming.
  • Keep a mix of offline evaluations, staged evaluations, and production monitoring.
  • Make sure evaluation results predict production behavior.

Set the quality bar

  • Define release criteria for agent changes. Approve or block releases with data.
  • Write clear documentation for the evaluation methods.
  • Report results to engineering and to leadership.

What We Require

AI and agent evaluation (4+ years)

  • 4 or more years of hands-on work in evaluation of LLM systems and agents.
  • Hands-on skill with RAGAs, Promptfoo, and DeepEval. Skill with Braintrust, LangSmith, W&B Weave, or OpenAI Evals is a plus.
  • Proof that you can create KPIs and metrics from zero: select the metrics, set the baselines, set the thresholds, and get agreement from stakeholders.
  • Proof that you can standardize evaluation across teams: shared metric definitions, reusable templates, versioned datasets, and one report format.
  • Experience with evaluation in more than one modality: text plus image, audio, video, or structured documents.
  • Knowledge of the agent evaluation problem space: trajectory evaluation, outcome evaluation, tool-use correctness, multi-turn behavior, non-determinism, and prompt sensitivity.
  • Experience with LLM-as-judge design: rubric design, calibration against human labels, and bias control.
  • Knowledge of statistics for evaluation: sampling, significance tests, inter-rater agreement, and variance across runs.
  • Strong Python skills. Experience with test harnesses, data pipelines, or evaluation tools that you built.
  • Skill with tracing and observability for agent systems. Example: OpenTelemetry-style traces and structured logs of tool calls.

Core software QA

  • Knowledge of QA fundamentals: test plans, test-case design, functional tests, regression tests, integration tests, and defect management.
  • Hands-on skill with test automation frameworks. Examples: pytest, Playwright, Selenium.
  • Hands-on skill with API tests. Examples: Postman, REST-assured.
  • Experience with test suites in CI/CD pipelines. Examples: GitHub Actions, Jenkins, GitLab CI.
  • Skill with bug tracking and test management tools. Examples: Jira, TestRail, Xray.
  • Ability to write clear defect reports that engineers can reproduce.
  • Knowledge of non-functional tests: performance, load, and reliability.

What Is a Plus

  • Experience with red-teaming, safety evaluations, or adversarial tests of AI systems.
  • Experience with human annotation work: labeling guidelines, calibration sessions, and inter-annotator agreement.
  • Background in CI/CD, infrastructure-as-code, or platform engineering.
  • Contributions to open-source evaluation tools. Published work on evaluation methods.

What Success Looks Like

  • 90 days: You know the current agent harness. You found the gaps in the current evaluations. A baseline regression suite runs on each change.
  • 6 months: Evaluation results are trusted release gates. Failure types are tracked and decrease. Engineers come to you before they ship.
  • 12 months: Other teams use the evaluation platform on their own. Offline results predict production quality. Our quality bar is a competitive advantage.

Why This Role Matters

Agents fail in ways that normal software does not. They fail silently. They fail at random. They fail on step ten of a long task. Users trust an AI product only when the evaluation is rigorous. You own that rigor.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Quality Engineer
AI Quality Engineer

Allegis Group Services, Inc. • India

On-site
INR 1,200,000 - 2,400,000
Senior AI Evaluation & Reliability Engineer
Senior AI Evaluation & Reliability Engineer

Aubergine Solutions Pvt. Ltd. • Ahmedabad District

On-site
INR 3,000,000 - 6,000,000
Great Place To Work certified
67503 Agent Evaluation & Instrumentation Engineer
67503 Agent Evaluation & Instrumentation Engineer

Cephas Consultancy Services Private Limited • Pune District

On-site
INR 3,000,000 - 4,500,000
Quality Assurance Engineer
Quality Assurance Engineer

Valiance Solutions • Dadri

On-site
INR 800,000 - 1,500,000
AI Agent Evaluation Engineer
AI Agent Evaluation Engineer

Letitbex AI • Telangana

On-site
INR 4,000,000 - 7,500,000
Agent Evaluation & Instrumentation Engineer
Agent Evaluation & Instrumentation Engineer

NCS Group • Pune District

On-site
INR 1,800,000 - 3,000,000
AI QA Engineer
AI QA Engineer

Huptech Hr Solutions • Ahmedabad District

On-site
INR 1,200,000 - 2,400,000
Agentic QA
Agentic QA

Horizon Industries International Limited • Dadri

On-site
INR 4,000,000 - 7,000,000
QA Lead- AI Evaluation & Quality
QA Lead- AI Evaluation & Quality

National e-Governance Division • New Delhi

On-site
INR 1,500,000 - 1,900,000
AI Engineer
AI Engineer

Qentelli • Hyderabad

On-site
INR 4,000,000 - 7,000,000