Senior AI Test Engineer

Alternative Path

Gurugram District

On-site

INR 800,000 - 1,400,000

Full time

32 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Alternative Path is seeking a Senior AI Test Engineer to lead QA for a client engagement in the financial data space. You will own the test strategy for AI agent systems, including chatbots, RAG-based research and tool-calling agents, ensuring accuracy, reliability, and safety before end-users access financial data.

You will design eval suites, curate golden datasets, and coordinate with the client's engineering team to mitigate risk.

Qualifications

  • 3+ years in software QA/testing with hands-on senior-level experience evaluating LLM-based or agentic AI systems
  • Direct experience with eval frameworks/tools including LangSmith or equivalents
  • Strong understanding of RAG architecture and evaluation of retrieval vs generation quality
  • Practical familiarity with MCP (Model Context Protocol) or similar agent tool-calling frameworks
  • Strong Python skills for building test harnesses, scoring pipelines, and automation
  • Experience designing evaluation approaches for non-deterministic systems
  • Comfortable working directly with client engineering teams and communicating risks
  • Client-facing experience in a consulting QA environment (not internal product)

Responsibilities

  • Own the test and evaluation strategy for the client's AI agent systems
  • Design and build eval suites covering accuracy, groundedness, task completion, latency, and safety
  • Build and curate golden datasets reflecting real client use cases and financial-data edge cases
  • Use eval/observability tools to trace, score, and monitor agent runs
  • Define and track RAG-specific metrics including retrieval precision/recall and citation accuracy
  • Evaluate MCP-based tool-calling behavior including correct tool selection and multi-step execution
  • Design regression testing to protect against degradation from changes
  • Run red-teaming for hallucinations, prompt injection, and unsafe outputs
  • Act as the primary quality contact with client stakeholders, translating findings into risk assessments
  • Mentor junior QA resources and scale engagement
  • Build automation to integrate evals into CI pipelines

Skills

Python
LLM evaluation
QA testing
Client communication
Non-deterministic systems

Tools

LangSmith
Ragas
DeepEval
PromptFoo
Braintrust
Langfuse

Job description

Alternative Path is looking for a Senior AI Test Engineer to lead quality assurance on a client engagement with a leading financial data and analytics company that has built AI agent systems (chatbots, RAG-based research/report generation, MCP-based tool-calling agents) on top of its proprietary datasets for its own clients. This is a senior, client-facing role: you'll own the test strategy for these systems end-to-end, working directly with the client's engineering team to validate accuracy, reliability, and safety before these agents reach financial industry end-users — where hallucinations or bad data retrieval carry real business risk.

This is not conventional QA. You'll be building evaluation frameworks and datasets, running systematic evals, and defining quality bars for non-deterministic, LLM-driven systems operating over complex financial data.

About Alternative Path

Alternative Path is a strategic operations consulting firm, founded in 2020, that helps growth-stage and enterprise organizations build high-performing, India-based teams across Technology Ops, Data Ops, Business Intelligence, and Finance. Every engagement is senior-led, and the firm has delivered 50+ engagements for clients across the US, Europe, and Australia.

What You'll Do
  • Own the test and evaluation strategy for the client's AI agent systems — chatbots, RAG-based research/report generation, and MCP-based tool-calling agents
  • Design and build eval suites covering accuracy, groundedness/faithfulness to source data, task completion, latency, and safety
  • Build and curate golden datasets reflecting real client use cases and financial-data edge cases (ambiguous queries, stale data, conflicting sources, numerical precision)
  • Use eval/observability tooling such as LangSmith (or equivalents — Ragas, DeepEval, Braintrust, Langfuse, PromptFoo) to trace, score, and monitor agent runs
  • Define and track RAG-specific metrics (retrieval precision/recall, hallucination rate, citation/source accuracy) — critical given the systems sit on top of financial datasets where correctness is non-negotiable
  • Evaluate MCP-based tool-calling behavior — correct tool selection, multi-step task execution, failure handling
  • Design regression testing so prompt, model, or data-pipeline changes don't silently degrade quality
  • Run structured red-teaming for hallucinations, prompt injection, and unsafe/incorrect financial outputs
  • Act as the primary quality point of contact with the client's engineering stakeholders — translating findings into clear, actionable risk assessments
  • Mentor/guide any junior QA resources staffed onto the engagement as it scales
  • Build automation to integrate evals into CI pipelines where feasible
What We're Looking For
Must-haves
  • 3+ years in software QA/testing, with hands-on senior-level experience evaluating LLM-based or agentic AI systems (not just traditional QA)
  • Direct experience with eval frameworks/tools — LangSmith strongly preferred; also acceptable: Ragas, DeepEval, PromptFoo, Braintrust, Langfuse
  • Strong understanding of RAG architecture and how to independently evaluate retrieval quality vs. generation quality
  • Practical familiarity with MCP (Model Context Protocol) or comparable agent tool-calling frameworks
  • Strong Python skills for building test harnesses, scoring pipelines, and automation
  • Experience designing evaluation approaches for non-deterministic systems (probabilistic scoring, not binary pass/fail)
  • Comfortable working directly with client engineering teams — strong communication, able to explain quality risk to both technical and business stakeholders
  • Prior experience in a client services / consulting QA environment (this is a client-facing engagement, not an internal product team)
Nice-to-haves
  • Experience testing AI systems in fintech, capital markets, or data/analytics domains
  • Familiarity with financial data accuracy/compliance considerations (numerical precision, source traceability, audit trails)
  • Exposure to LLM red-teaming/safety evaluation practices
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Test Engineer
Senior AI Test Engineer

Alternative Path • India

Remote
INR 1,500,000 - 3,000,000
Senior Quality Assurance Engineer
Senior Quality Assurance Engineer

Alternative Path • India

On-site
INR 900,000 - 1,200,000
AI QA Engineer
AI QA Engineer

Huptech Hr Solutions • Ahmedabad District

On-site
INR 1,200,000 - 2,400,000
QA Engineer
QA Engineer

YO IT Consulting • Mumbai

On-site
INR 800,000 - 1,400,000
AI Engineer
AI Engineer

Qentelli • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Quality Assurance Engineer
Quality Assurance Engineer

Meril • Gujarat

On-site
INR 1,000,000 - 1,500,000
Opportunity to work with Generative AI
Technical environment treating QA as engineering
Freedom to implement new testing methodologies
Principal Software Engineer
Principal Software Engineer

Cadence Design Systems • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Quality Assurance AI Engineer II CloudAngles
Senior Quality Assurance AI Engineer II CloudAngles

Cloud Angles Digital Transformation • Hyderabad

Hybrid
INR 1,200,000 - 1,800,000
Principal Software Engineer
Principal Software Engineer

Cadence • Bengaluru

On-site
INR 4,500,000 - 7,500,000
AI Senior QA
AI Senior QA

Hiver • Bengaluru

On-site
INR 1,500,000 - 2,200,000