Get more replies from employers
Send a job-specific resume in minutes.
Fermi AI is seeking an AI Evaluation Engineer for on-site work in Bengaluru, focused on building evaluation scaffolding and quality layers for AI features. You will manage eval harnesses, test infrastructure, and release gates across AI assistants like Claude and ChatGPT, adapting to host differences and evolving models.
The role requires 3–6 years of experience, strong Python or TypeScript skills, and hands-on API testing with Playwright or Cypress.
Building with AI is easy to prototype, but proving reliability in production is a major challenge. In AI-native codebases, verification is the key bottleneck for scaling capabilities. As an AI Evaluation Engineer, you will take ownership of the evaluation scaffolding and quality layer, including eval harnesses for AI-facing features, test infrastructure, and release gates. You’ll play a critical role in monitoring how our product behaves across AI assistants (e.g., Claude, ChatGPT), accounting for differences by host and continual changes.
Location: HSR, Bengaluru (On-site)
Experience: 3-6 years
Building with AI is easy to prototype, but proving reliability in production is a major challenge. In AI-native codebases, verification is the key bottleneck for scaling capabilities. As an AI Evaluation Engineer, you will take ownership of the evaluation scaffolding and quality layer, including eval harnesses for AI-facing features, test infrastructure, and release gates. You’ll play a critical role in monitoring how our product behaves across AI assistants (e.g., Claude, ChatGPT), accounting for differences by host and continual changes.
Background in SDET/QA automation or ML evaluation, with proven ownership of test or evaluation infrastructure—not just executing tests
Strong programming skills in Python or TypeScript, with experience in API-level testing and tools like Playwright or Cypress
Familiarity with LLM applications or a demonstrated interest in evaluating non-deterministic systems
Highly autonomous and able to define your own workflows and processes for evaluation and verification