Senior AI Evaluation & Reliability Engineer

Aubergine Solutions Pvt. Ltd.

Ahmedabad District

On-site

INR 3,000,000 - 6,000,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Great Place To Work certified

Job summary

Aubergine Solutions Pvt. Ltd.

in Ahmedabad, India, invites a Senior AI Evaluation & Reliability Engineer to design production-grade evaluation systems for LLMs, RAG apps, and multi-agent systems, aligning AI performance with business value. You will mentor engineers, shape evaluation frameworks, and collaborate with North American enterprise clients to deliver trusted AI solutions, ROI dashboards, and scalable AI quality processes.

Qualifications

  • Hands-on experience with LLM evaluation, calibration, and bias.
  • Strong Python and/or TypeScript programming skills.
  • Experience building production-grade evaluation pipelines.
  • Familiarity with CI/CD and automated testing.
  • Knowledge of vector databases and retrieval architectures.
  • Experience with enterprise client projects and ROI framing.

Responsibilities

  • Design evaluation strategies linking AI performance to business outcomes.
  • Build production-grade evaluation pipelines for LLMs, RAG, and agents.
  • Establish benchmarks, risk thresholds, and quality scorecards.
  • Mentor engineers and lead knowledge-sharing sessions.
  • Collaborate with North American clients to translate metrics into business decisions.
  • Optimize evaluation cost, latency, and throughput.

Skills

Python
TypeScript
Backend APIs
CI/CD
Vector DBs
AI Evaluation
LLM Evaluation
RAG Evaluation
Calibration
Stakeholder communication

Tools

DeepEval
Rugas
TruLens
LangSmith
DSPy
Promptfoo
Pinecone
Qdrant
Milvus
Chroma

Job description

Senior AI Evaluation & Reliability Engineer

8 - 10 Years

Full-Time

Why Aubergine

Aubergine is a global transformation and innovation partner , shaping next-gen digital products through consulting-informed execution that integrates strategy, design, and development.

Since 2013, we’ve built 400+ B2B and B2C products worldwide , turning powerful ideas into impact-driven experiences. We are one of the top global B2B companies on Clutch , rated highest among more than 80,000 technology service providers.

With more than 150 digital thinkers , we are home to some of the brightest, most passionate people around the world who are committed to delivering excellence.

We’re not just another workplace. We’re officially Great Place To Work® certified , with an exceptional trust index rating, making Aubergine a community where you can thrive and grow.

Role Overview: Build AI Systems We Can Trust

We are looking for a Senior AI Evaluation & Reliability Engineer who is passionate about solving one of the most important challenges in AI:

How do we know an AI system is actually working, improving, and delivering business value?

You will design and build production-grade evaluation systems for LLMs, RAG applications, and multi-agent systems, while helping enterprise clients and our engineering teams adopt AI with confidence.

This role goes beyond building evals. You will be a trusted AI consultant to clients, a technical mentor to engineers, and a key contributor to our journey towards becoming an AI superagency.

What You Will Own
  • Make AI Performance & ROI Measurable
    • Build evaluation strategies that connect AI performance to business outcomes and ROI.
    • Define quality benchmarks, SLOs, risk thresholds, and success criteria.
    • Track hallucinations, reliability issues, and production risks.
    • Build executive-friendly AI quality and ROI scorecards.
    • Help clients determine where AI should be autonomous, supervised, or avoided.
    • Don't just measure model accuracy. Measure business impact.
  • Build Production-Grade Evaluation Pipelines
    • Architect automated evaluation pipelines for LLMs, RAG, and agentic systems.
    • Integrate evaluations into CI/CD using GitHub Actions, GitLab CI, or equivalent.
    • Build regression suites for prompts, models, tools, and workflows.
    • Measure faithfulness, context precision, answer relevance, semantic drift, and task completion.
    • Establish statistically meaningful benchmarks and continuously monitor AI quality.
    • Evaluation should become part of engineering, not a final QA step.
    • Design reference-based and reference-free LLM evaluation frameworks.
    • Create structured rubrics and scoring systems.
    • Build calibration loops using human-labelled datasets.
    • Identify and mitigate judge biases such as position, verbosity, and self-preference bias.
    • Measure judge reliability and optimize evaluation quality, latency, and cost.
  • Own Evaluation Economics
    • LLM evaluations can become expensive quickly.
    • You will:
      • Design tiered evaluation strategies.
      • Use deterministic and heuristic graders for simple checks.
      • Reserve powerful LLM judges for complex evaluations.
      • Optimise batching, concurrency, and parallel execution.
      • Track evaluation cost and balance quality, latency, and inference spend.
  • Benchmark RAG & Agentic Systems
    • Build systematic benchmarks across:
    • Retrieval quality, faithfulness, and answer relevance
    • Embedding, chunking and re-ranking strategies
    • Vector databases and retrieval architectures
    • Agent task completion and tool-calling accuracy
    • State retention, multi-step workflows and failure recovery
    • Latency, reliability and cost
    • Build synthetic datasets, golden test suites, and adversarial scenarios to continuously expand evaluation coverage.
Be a Trusted AI Consultant
  • You will work directly with North American enterprise clients as a technical AI advisor.
  • You will:
    • Lead AI architecture and evaluation discussions.
    • Translate complex AI metrics into clear business recommendations.
    • Define AI quality, reliability, and risk frameworks.
    • Present evaluation telemetry and ROI scorecards to technical and business stakeholders.
    • Challenge assumptions and recommend the right AI solution, even when that means saying "don't use AI here."
    • You should be equally comfortable discussing LLM evaluation with engineers and ROI with a CTO/COO/CEO/CFO.

You won't just help us deliver AI solutions. You'll help us win the right AI problems to solve.

Coach Engineers. Raise the Bar.

As our organisation evolves towards an AI superagency, you will help shape how we build AI.

You will:

  • Mentor engineers working on AI and LLM systems.
  • Establish AI engineering and evaluation best practices.
  • Conduct technical workshops and knowledge-sharing sessions.
  • Review AI architectures and evaluation strategies.
  • Help engineers move from prompt experimentation to disciplined AI engineering.
  • Build reusable frameworks, playbooks, and internal accelerators.
  • Stay ahead of emerging AI technologies and bring valuable ideas into the organisation.
What We Are Looking For
AI Evaluation & Reliability
  • Strong practical experience with LLM-as-a-Judge.
  • Experience with calibration, human ground truth, and evaluation bias.
  • Strong understanding of RAG and agent evaluation.
  • Knowledge of Faithfulness, Context Precision, Answer Relevance, Task Completion, and semantic similarity.
  • Experience with DeepEval, Ragas, TruLens, LangSmith, DSPy, Promptfoo, or equivalent.
  • Advanced Python and/or TypeScript.
  • Strong backend, API, and asynchronous systems experience.
  • Experience with CI/CD and automated testing.
  • Experience with vector databases such as Pinecone, Qdrant, Milvus, or Chroma.
  • Understanding of LLM infrastructure, observability, and inference economics.
  • Experience working directly with enterprise clients.
  • Excellent communication and presentation skills.
  • Strong consulting and problem-solving mindset.
  • Ability to translate technical complexity into business outcomes.
  • Strong mentoring and coaching abilities.
  • Curiosity, ownership, and enthusiasm for emerging AI technologies.
Why This Role Matters

The next generation of AI companies won't simply be the ones with access to the best models.

They will be the ones that can measure AI, trust AI, improve AI, and turn AI into business value.

If you want to build that future with us, we'd love to hear from you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Quality Assurance Engineer
AI Quality Assurance Engineer

Thirdeye Data Inc. • Bengaluru

Hybrid
INR 1,200,000 - 2,000,000
Lead AI Engineer - Agentic Engineering
Lead AI Engineer - Agentic Engineering

Blend360 India • Hyderabad

Hybrid
INR 3,000,000 - 5,000,000
Senior AI / LLM Engineer
Senior AI / LLM Engineer

ThinkScoop Inc • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Competitive salary with performance-linked bonuses
Remote-friendly with async-first culture
Access to latest LLM platforms and tooling budgets
+1
Quality Assurance Engineer
Quality Assurance Engineer

Valiance Solutions • Dadri

On-site
INR 800,000 - 1,500,000
Senior AI Engineer - India
Senior AI Engineer - India

Acrotrend - A NowVertical Company • Navi Mumbai

On-site
INR 3,000,000 - 5,400,000
Competitive salary
Global client exposure
Forward Deployed Engineer
Forward Deployed Engineer

Insight Global • Hyderabad

On-site
INR 3,500,000 - 6,500,000
Newpage - Full stack AI engineer - Python/React.js
Newpage - Full stack AI engineer - Python/React.js

Newpage Solutions • Maharashtra

On-site
INR 2,400,000 - 4,200,000
Principal AI Engineer
Principal AI Engineer

WillWare Technologies • Bengaluru

On-site
INR 4,000,000 - 6,500,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Insaito Software • India

On-site
INR 1,200,000 - 1,800,000
Senior AI Engineer - Agentic
Senior AI Engineer - Agentic

Srijan Technologies PVT LTD • Gurugram District

On-site
INR 1,200,000 - 2,000,000