Stand out for this role — generate a tailored resume and cover letter in about a minute.
Meraki Labs in Bengaluru is seeking an AI Evaluation Engineer to own evaluation scaffolding and quality layers for live Proofline deployments, ensuring reliability across AI assistants.
You will build test infrastructure, eval harnesses, and release gates, with a focus on API-level testing using Python or TypeScript and Playwright or Cypress.
Work in a small, high-trust team where design reviews are thorough and automation is valued, while releasing features daily in production.
Meraki Labs (founded by Mukesh Bansal & Peeyush Ranjan) builds and rapidly scales AI-first, "moonshot" startups. We're looking for a high-velocity, production-grade engineer to build 0-to-1 products alongside founders.
Proofline is live in production with institutional pilots underway. Building with AI is easy to prototype, but proving reliability in production is a major challenge. In AI-native codebases, verification is the key bottleneck for scaling capabilities. As an AI Evaluation Engineer, you will take ownership of the evaluation scaffolding and quality layer, including eval harnesses for AI-facing features, test infrastructure, and release gates. You’ll play a critical role in monitoring how our product behaves across AI assistants (e.g., Claude, ChatGPT), accounting for differences by host and continual changes.
Small team, high trust, written decisions. Designs get adversarial review before code; PRs get automated review driven to zero open findings; features aren't done until verified on a live system. AI agents do a large share of the implementation—your leverage is judgment: framing the problem, freezing the right design, and knowing when the machine is wrong.