AI Evaluation & Test Engineer

BharatGen

Mumbai

On-site

INR 1,500,000 - 2,500,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

A leading tech company in Mumbai is seeking an experienced AI Evaluation & Test Engineer to join their team. The ideal candidate will build and maintain AI evaluation pipelines, ensuring that generative AI models operate effectively and safely. Applicants should have a strong background in software testing, experience with AI/ML testing, and proficiency in relevant tools and programming languages. This role offers a dynamic environment that encourages collaboration and innovation in AI applications.

Qualifications

  • 5+ years in manual and automation testing of software products.
  • At least 2 years in evaluating and testing AI/ML products.
  • Strong skills in writing test plans and executing test cases.

Responsibilities

  • Build and maintain AI evaluation pipelines.
  • Implement evaluation and testing automation for AI systems.
  • Define and implement criteria for release gates in CI/CD pipeline.
  • Collaborate with cross-functional teams to enhance AI user experience.

Skills

Software testing fundamentals
Analytical skills
Attention to detail
Collaboration
Go-getter attitude
Python proficiency
Experience with testing frameworks

Education

Bachelor’s or Master’s degree in CS/CE/IT/EE/E&TC

Tools

Pytest
Selenium
Robot Framework
AI evaluation frameworks

Job description

We are looking for an AI Evaluation & Test Engineer to join our growing team to ensure that our generative AI models and applications are safe, accurate, trustworthy, and deliver an elegant user experience. You will serve as the first customer of our AI systems. This role is ideal for product-minded engineers who obsess over product quality and customer-centricity, and are passionate about shaping the behavior of AI systems in the real world.

Key Responsibilities
  • Build and maintain AI evaluation pipelines to test, measure, and evaluate the behavior and performance of AI systems.
  • Implement traces, spans, and session tracking for observability and identify error propagation in multi-step pipelines.
  • Define AI quality metrics and KPIs around factuality, faithfulness, toxicity, grounding precision/recall, latency, cost, etc., with clear acceptance bars.
  • Implement evaluation and testing automation to enable end-to-end system and regression testing at scale.
  • Define criteria for and implement release gates in the CI/CD pipeline.
  • Define criteria for and implement release gates in the CI/CD pipeline.
  • Find creative ways to break products.
  • Assist in root cause analysis and troubleshooting of bugs and field issues.
  • Collaborate with cross-functional teammates from product, engineering, linguistics,, and customer support to shape human-AI interaction paradigms and ensure that our AI models and applications deliver the desired outcome and user experience.
  • Bachelor’s or Master’s degree in CS/CE/IT/EE/E&TC or related fields with 5+ years of experience in manual and automation testing of software products, with at least 2 years in evaluating and testing AI/ML products.
  • Strong software testing fundamentals and expertise in writing test plans, executing test cases, and generating detailed reports and dashboards.
  • Strong analytical and debugging skills, and attention to detail.
  • Proficiency in Python, scripting, and software testing automation frameworks and tools such as Pytest, Selenium, Robot Framework, etc.
  • Working knowledge of generative AI models, AI agents, and related concepts such as retrieval augmented generation (RAG), prompt engineering, context engineering, explainability, traceability, observability, guard rails, reasoning, specificity, etc.
  • Sound understanding of the fundamental differences in the approach for testing conventional software versus evaluating generative AI systems.
  • Team player with excellent interpersonal skills and the ability to collaborate effectively with remote and cross- functional team members.
  • Go-getter attitude and ability to flourish in a fast-paced, startup environment.
  • Experience in any of the following would be a big plus.
  • AI evaluation frameworks such as Arize, Braintrust, DeepEval, LangSmith, Ragas
  • AI safety and red teaming experience, e.g., prompt injection, jailbreak, adversarial and stress testing.
  • Different types of AI evaluation methods, e.g, Human-in-the-loop, LLM-as-a-Judge.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Forward Deployed Engineer - AI Assurance
Forward Deployed Engineer - AI Assurance

Systems Limited • India

On-site
INR 1,200,000 - 1,800,000
Functional AI Tester - GenAI
Functional AI Tester - GenAI

Michelin • Pune District

On-site
INR 1,200,000 - 1,800,000
Agentic AI Test Engineer
Agentic AI Test Engineer

Deutsche Telekom Digital Labs • Gurugram District

On-site
INR 1,620,000 - 1,980,000
QA Engineer with AI
QA Engineer with AI

Infoya Inc. • Navalur

On-site
INR 900,000 - 1,500,000
Hiring | Gen AI Quality Assurance (BLR/Pune)
Hiring | Gen AI Quality Assurance (BLR/Pune)

2COMS Consulting Pvt. Ltd. • Bengaluru

On-site
INR 450,000 - 750,000
Hiring || Gen AI Quality Assurance (BLR/Pune)
Hiring || Gen AI Quality Assurance (BLR/Pune)

2coms • Bengaluru

On-site
INR 2,500,000 - 4,000,000
AI QA & Data Quality Specialist (Chennai / Pune )
AI QA & Data Quality Specialist (Chennai / Pune )

Money Forward India • Chennai District

On-site
INR 1,400,000 - 2,100,000
Quality Intelligence Engineer _AI-Powered Quality Engineering & Agent
Quality Intelligence Engineer _AI-Powered Quality Engineering & Agent

Capgemini • Hyderabad

Hybrid
INR 1,400,000 - 2,200,000
AI Evaluation Engineer - Proofline
AI Evaluation Engineer - Proofline

Fermi AI • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Principal Software Engineer
Principal Software Engineer

Cadence • Bengaluru

On-site
INR 4,500,000 - 7,500,000