AI Evaluation Engineer - Proofline

Fermi AI

Bengaluru

On-site

INR 1,500,000 - 2,100,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Fermi AI is seeking an AI Evaluation Engineer for on-site work in Bengaluru, focused on building evaluation scaffolding and quality layers for AI features. You will manage eval harnesses, test infrastructure, and release gates across AI assistants like Claude and ChatGPT, adapting to host differences and evolving models.

The role requires 3–6 years of experience, strong Python or TypeScript skills, and hands-on API testing with Playwright or Cypress.

Qualifications

  • Background in SDET/QA automation or ML evaluation.
  • Strong programming skills in Python or TypeScript, with experience in API-level testing and tools like Playwright or Cypress.
  • Familiarity with LLM applications or evaluating non-deterministic systems.

Responsibilities

  • Build and maintain evaluation harnesses for AI-facing features to measure system quality.
  • Own end‑to‑end and API test infrastructure (Playwright‑class) for daily releases.
  • Design host‑behavior probes using scripted sessions across diverse AI assistants.
  • Gate production releases through thorough user acceptance testing and quality reporting.

Skills

SDET/QA automation
ML evaluation
Python
TypeScript

Tools

Playwright
Cypress
API testing

Job description

About The Role

Building with AI is easy to prototype, but proving reliability in production is a major challenge. In AI-native codebases, verification is the key bottleneck for scaling capabilities. As an AI Evaluation Engineer, you will take ownership of the evaluation scaffolding and quality layer, including eval harnesses for AI-facing features, test infrastructure, and release gates. You’ll play a critical role in monitoring how our product behaves across AI assistants (e.g., Claude, ChatGPT), accounting for differences by host and continual changes.


Location: HSR, Bengaluru (On-site)


Experience: 3-6 years


About The Role

Building with AI is easy to prototype, but proving reliability in production is a major challenge. In AI-native codebases, verification is the key bottleneck for scaling capabilities. As an AI Evaluation Engineer, you will take ownership of the evaluation scaffolding and quality layer, including eval harnesses for AI-facing features, test infrastructure, and release gates. You’ll play a critical role in monitoring how our product behaves across AI assistants (e.g., Claude, ChatGPT), accounting for differences by host and continual changes.


Responsibilities


  • Build and maintain evaluation harnesses for AI-facing features to measure and tune system quality (e.g., capture quality, retrieval quality, guidance quality)

  • Own end‑to‑end and API test infrastructure (Playwright‑class), supporting a continuous, daily‑release cycle

  • Design and execute host‑behavior probes using scripted user sessions across diverse AI assistants, ensuring product behavior aligns with expectations

  • Gate production releases through thorough user acceptance testing (UAT), regression analysis, and quality reporting


Requirements

Background in SDET/QA automation or ML evaluation, with proven ownership of test or evaluation infrastructure—not just executing tests


Strong programming skills in Python or TypeScript, with experience in API-level testing and tools like Playwright or Cypress


Familiarity with LLM applications or a demonstrated interest in evaluating non-deterministic systems


Highly autonomous and able to define your own workflows and processes for evaluation and verification

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation & Test Engineer
AI Evaluation & Test Engineer

BharatGen • Mumbai

On-site
INR 1,500,000 - 2,500,000
AI Senior QA
AI Senior QA

Hiver • Bengaluru

On-site
INR 1,500,000 - 2,200,000
Forward Deployed Engineer - AI Assurance
Forward Deployed Engineer - AI Assurance

Systems Limited • India

On-site
INR 1,200,000 - 1,800,000
Principal Software Engineer
Principal Software Engineer

Cadence • Bengaluru

On-site
INR 4,500,000 - 7,500,000
AI Senior QA
AI Senior QA

Hiver, Inc. • Bengaluru

On-site
INR 2,200,000 - 3,400,000
Senior Quality Assurance Engineer
Senior Quality Assurance Engineer

Ecolibrium • Bengaluru Urban

On-site
INR 1,200,000 - 2,400,000
AI QE Engineer - Manager
AI QE Engineer - Manager

PwC • Hyderabad, Chennai District, Bengaluru

On-site
INR 1,500,000 - 2,100,000
AI Evaluation Expert - Remote
AI Evaluation Expert - Remote

YO AI Labs • Bengaluru

Remote
INR 3,321,000 - 7,971,000
AI Evaluation Specialist - Remote
AI Evaluation Specialist - Remote

YO AI Labs • Hyderabad

Remote
INR 3,318,000 - 5,309,000
AI Evaluation Specialist - Remote
AI Evaluation Specialist - Remote

YO AI Labs • Dadri

Remote
INR 5,780,000 - 8,671,000