AI Evaluation Engineer

Meraki Labs

Bengaluru

On-site

INR 1,800,000 - 2,600,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Meraki Labs in Bengaluru is seeking an AI Evaluation Engineer to own evaluation scaffolding and quality layers for live Proofline deployments, ensuring reliability across AI assistants.

You will build test infrastructure, eval harnesses, and release gates, with a focus on API-level testing using Python or TypeScript and Playwright or Cypress.

Work in a small, high-trust team where design reviews are thorough and automation is valued, while releasing features daily in production.

Qualifications

  • 5+ years of experience in SDET/QA automation or ML evaluation.
  • Own test or evaluation infrastructure, not just executing tests.
  • Strong programming skills in Python or TypeScript, with API-level testing experience.

Responsibilities

  • Build and maintain evaluation harnesses for AI-facing features to measure and tune system quality.
  • Own end-to-end and API test infrastructure supporting a daily-release cycle.
  • Design and execute host-behavior probes across diverse AI assistants to ensure correct product behavior.
  • Gate production releases with thorough UAT, regression analysis, and quality reporting.

Skills

Python
TypeScript
API testing
ML evaluation
Automation workflows

Tools

Playwright
Cypress

Job description

Meraki Labs (founded by Mukesh Bansal & Peeyush Ranjan) builds and rapidly scales AI-first, "moonshot" startups. We're looking for a high-velocity, production-grade engineer to build 0-to-1 products alongside founders.


About the Role

Proofline is live in production with institutional pilots underway. Building with AI is easy to prototype, but proving reliability in production is a major challenge. In AI-native codebases, verification is the key bottleneck for scaling capabilities. As an AI Evaluation Engineer, you will take ownership of the evaluation scaffolding and quality layer, including eval harnesses for AI-facing features, test infrastructure, and release gates. You’ll play a critical role in monitoring how our product behaves across AI assistants (e.g., Claude, ChatGPT), accounting for differences by host and continual changes.


Responsibilities


  • Build and maintain evaluation harnesses for AI-facing features to measure and tune system quality (e.g., capture quality, retrieval quality, guidance quality)

  • Own end-to-end and API test infrastructure (Playwright-class), supporting a continuous, daily-release cycle

  • Design and execute host-behavior probes using scripted user sessions across diverse AI assistants, ensuring product behavior aligns with expectations

  • Gate production releases through thorough user acceptance testing (UAT), regression analysis, and quality reporting


Requirements


  • Experience - 5+ years

  • Background in SDET/QA automation or ML evaluation, with proven ownership of test or evaluation infrastructure and not just executing tests

  • Strong programming skills in Python or TypeScript, with experience in API-level testing and tools like Playwright or Cypress

  • Familiarity with LLM applications or a demonstrated interest in evaluating non-deterministic systems

  • Highly autonomous and able to define your own workflows and processes for evaluation and verification


How We Work

Small team, high trust, written decisions. Designs get adversarial review before code; PRs get automated review driven to zero open findings; features aren't done until verified on a live system. AI agents do a large share of the implementation—your leverage is judgment: framing the problem, freezing the right design, and knowing when the machine is wrong.


You Should Apply If


  • You enjoy turning messy problems into concrete solutions.

  • You thrive in fast-moving, low-process environments.

  • You're excited about production engineering (monitoring, reliability, cost, latency).


You Should Not Apply If


  • You want remote/hybrid (this is onsite Bangalore only).

  • You prefer narrow tickets and minimal ambiguity.

  • You prefer highly structured, slow-moving product organizations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer - Proofline
AI Evaluation Engineer - Proofline

Fermi AI • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Chief Technology Officer
Chief Technology Officer

Meraki Labs • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Chief Technology Officer - Proofline
Chief Technology Officer - Proofline

Fermi AI • Bengaluru

On-site
INR 6,000,000 - 9,000,000
On-site in Bengaluru
Competitive compensation
Performance incentives
Founding AI Engineer
Founding AI Engineer

Meraki Labs • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Evals Engineer
Evals Engineer

PingAura AI Technologies Private Limited. • Mumbai

On-site
INR 1,000,000 - 1,500,000
Founding-tier ownership
Direct collaboration with AI experts
Competitive salary
Chief Technology Officer - Proofline
Chief Technology Officer - Proofline

Meraki-Labs • Bengaluru

On-site
INR 2,400,000 - 4,200,000
Senior AI Evaluation & Reliability Engineer
Senior AI Evaluation & Reliability Engineer

Aubergine Solutions Pvt. Ltd. • Ahmedabad District

On-site
INR 3,000,000 - 6,000,000
Great Place To Work certified
Applied AI Engineer
Applied AI Engineer

PingAura AI Technologies Private Limited. • Mumbai

On-site
INR 1,000,000 - 1,500,000
AI Agent Engineer
AI Agent Engineer

PingAura AI Technologies Private Limited. • Mumbai

On-site
INR 1,200,000 - 2,000,000
Founding Engineer - Olympiz
Founding Engineer - Olympiz

Meraki Labs • Bengaluru

On-site
INR 2,500,000 - 4,000,000