Senior AI Evaluation & Reliability Engineer
8 - 10 Years
Full-Time
Why Aubergine
Aubergine is a global transformation and innovation partner , shaping next-gen digital products through consulting-informed execution that integrates strategy, design, and development.
Since 2013, we’ve built 400+ B2B and B2C products worldwide , turning powerful ideas into impact-driven experiences. We are one of the top global B2B companies on Clutch , rated highest among more than 80,000 technology service providers.
With more than 150 digital thinkers , we are home to some of the brightest, most passionate people around the world who are committed to delivering excellence.
We’re not just another workplace. We’re officially Great Place To Work® certified , with an exceptional trust index rating, making Aubergine a community where you can thrive and grow.
Role Overview: Build AI Systems We Can Trust
We are looking for a Senior AI Evaluation & Reliability Engineer who is passionate about solving one of the most important challenges in AI:
How do we know an AI system is actually working, improving, and delivering business value?
You will design and build production-grade evaluation systems for LLMs, RAG applications, and multi-agent systems, while helping enterprise clients and our engineering teams adopt AI with confidence.
This role goes beyond building evals. You will be a trusted AI consultant to clients, a technical mentor to engineers, and a key contributor to our journey towards becoming an AI superagency.
What You Will Own
- Make AI Performance & ROI Measurable
- Build evaluation strategies that connect AI performance to business outcomes and ROI.
- Define quality benchmarks, SLOs, risk thresholds, and success criteria.
- Track hallucinations, reliability issues, and production risks.
- Build executive-friendly AI quality and ROI scorecards.
- Help clients determine where AI should be autonomous, supervised, or avoided.
- Don't just measure model accuracy. Measure business impact.
- Build Production-Grade Evaluation Pipelines
- Architect automated evaluation pipelines for LLMs, RAG, and agentic systems.
- Integrate evaluations into CI/CD using GitHub Actions, GitLab CI, or equivalent.
- Build regression suites for prompts, models, tools, and workflows.
- Measure faithfulness, context precision, answer relevance, semantic drift, and task completion.
- Establish statistically meaningful benchmarks and continuously monitor AI quality.
- Evaluation should become part of engineering, not a final QA step.
- Design reference-based and reference-free LLM evaluation frameworks.
- Create structured rubrics and scoring systems.
- Build calibration loops using human-labelled datasets.
- Identify and mitigate judge biases such as position, verbosity, and self-preference bias.
- Measure judge reliability and optimize evaluation quality, latency, and cost.
- Own Evaluation Economics
- LLM evaluations can become expensive quickly.
- You will:
- Design tiered evaluation strategies.
- Use deterministic and heuristic graders for simple checks.
- Reserve powerful LLM judges for complex evaluations.
- Optimise batching, concurrency, and parallel execution.
- Track evaluation cost and balance quality, latency, and inference spend.
- Benchmark RAG & Agentic Systems
- Build systematic benchmarks across:
- Retrieval quality, faithfulness, and answer relevance
- Embedding, chunking and re-ranking strategies
- Vector databases and retrieval architectures
- Agent task completion and tool-calling accuracy
- State retention, multi-step workflows and failure recovery
- Latency, reliability and cost
- Build synthetic datasets, golden test suites, and adversarial scenarios to continuously expand evaluation coverage.
Be a Trusted AI Consultant
- You will work directly with North American enterprise clients as a technical AI advisor.
- You will:
- Lead AI architecture and evaluation discussions.
- Translate complex AI metrics into clear business recommendations.
- Define AI quality, reliability, and risk frameworks.
- Present evaluation telemetry and ROI scorecards to technical and business stakeholders.
- Challenge assumptions and recommend the right AI solution, even when that means saying "don't use AI here."
- You should be equally comfortable discussing LLM evaluation with engineers and ROI with a CTO/COO/CEO/CFO.
You won't just help us deliver AI solutions. You'll help us win the right AI problems to solve.
Coach Engineers. Raise the Bar.
As our organisation evolves towards an AI superagency, you will help shape how we build AI.
You will:
- Mentor engineers working on AI and LLM systems.
- Establish AI engineering and evaluation best practices.
- Conduct technical workshops and knowledge-sharing sessions.
- Review AI architectures and evaluation strategies.
- Help engineers move from prompt experimentation to disciplined AI engineering.
- Build reusable frameworks, playbooks, and internal accelerators.
- Stay ahead of emerging AI technologies and bring valuable ideas into the organisation.
What We Are Looking For
AI Evaluation & Reliability
- Strong practical experience with LLM-as-a-Judge.
- Experience with calibration, human ground truth, and evaluation bias.
- Strong understanding of RAG and agent evaluation.
- Knowledge of Faithfulness, Context Precision, Answer Relevance, Task Completion, and semantic similarity.
- Experience with DeepEval, Ragas, TruLens, LangSmith, DSPy, Promptfoo, or equivalent.
- Advanced Python and/or TypeScript.
- Strong backend, API, and asynchronous systems experience.
- Experience with CI/CD and automated testing.
- Experience with vector databases such as Pinecone, Qdrant, Milvus, or Chroma.
- Understanding of LLM infrastructure, observability, and inference economics.
- Experience working directly with enterprise clients.
- Excellent communication and presentation skills.
- Strong consulting and problem-solving mindset.
- Ability to translate technical complexity into business outcomes.
- Strong mentoring and coaching abilities.
- Curiosity, ownership, and enthusiasm for emerging AI technologies.
Why This Role Matters
The next generation of AI companies won't simply be the ones with access to the best models.
They will be the ones that can measure AI, trust AI, improve AI, and turn AI into business value.
If you want to build that future with us, we'd love to hear from you.