LLM / Agentic Evaluation Rig Engineer

PHIZENIX

Hyderabad

Hybrid

INR 1,500,000 - 2,000,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

PHIZENIX is seeking an LLM/Evaluation Engineer to build the system gating AI outputs for production readiness. You will own datasets, scorers, harnesses, and CI gates ensuring grounding and faithfulness before release.

Your role emphasizes evidence over vibes, multi-step agentic evaluation, and collaboration with the staff AI team to drive quality improvements across releases.

Qualifications

  • 4+ years in software/ML engineering building evaluation tools.
  • Experience measuring grounding, faithfulness, hallucination.
  • Strong Python and engineering practices.
  • Familiar with LLM eval frameworks.
  • FinTech or high-stakes AI domain a plus.

Responsibilities

  • Design evaluation datasets and ground-truth labels.
  • Build scorers for grounding, faithfulness, and hallucination.
  • Develop CI gates blocking regressive releases.
  • Collaborate with teams to integrate evaluation into release processes.

Skills

Python
ML Engineering
CI/CD
LLM Eval
Grounding & Faithfulness
Statistics / Metrics

Tools

LangSmith
GitHub Actions
LLM APIs

Job description

We are looking for an LLM / Agentic Evaluation Rig Engineer to build the system that decides whether our AI output is good enough to ship. Because our commentary sits next to externally reported financials, we cannot rely on vibes — grounding, faithfulness, and hallucination have to be measured, tracked, and gated before anything reaches a customer.

You own the evaluation infrastructure: the datasets, the scorers, the harnesses, and the CI gates that hold the AI and agentic layers to a hard quality bar. You are the team's source of truth on whether a model, prompt, or agent change is actually an improvement — and the one who blocks it if it isn't.

What makes this role different You define "good enough to ship" — your gates block regressions in grounding and faithfulness from reaching production. Evidence over vibes — every claim is checked against the verified source data it must be grounded in. Agentic evaluation — you evaluate multi-step reasoning flows, not just single prompts. Real leverage — your rig is how the whole AI team moves fast without breaking trust.

Responsibilities
Datasets & Scorers (35%)

Build and curate evaluation datasets, including adversarial and edge-case sets with ground-truth labels Build scorers for grounding, faithfulness, hallucination, factual consistency, and structured-output validity Combine rule-based checks, reference-based metrics, and LLM-as-judge where appropriate Verify generated claims map to verified source data — no unsupported statements

Harnesses & CI Gates (30%)

Build harnesses that run evaluations reproducibly across model, prompt, and agent versions Wire evaluation into CI so grounding / faithfulness regressions block releases Track quality over time with dashboards and clear pass / fail thresholds

Agentic Evaluation (25%)

Evaluate multi-step / agentic flows — routing, tool-use, verification, confirmation Build trace capture and step-level scoring for agent runs Detect where a flow silently degrades

Collaboration (10%)

Partner with the Staff AI Engineer to turn findings into model / prompt / orchestration improvements Partner with QA to integrate AI evaluation into the broader release process

Technical Stack
  • Evaluation LLM eval frameworks (promptfoo, DeepEval, Ragas, LangSmith)
  • LLM-as-judge, reference-based metrics
  • Dataset / ground-truth curation
  • AI & Orchestration
  • LLM APIs & managed LLMs (Bedrock / Vertex / Azure OpenAI)
  • RAG & agentic patterns (LangGraph)
  • Structured-output validation
  • Engineering Python CI/CD (GitHub Actions)
  • Dashboards & metrics tracking
What You’ll Build in Year One
  • A labeled evaluation dataset suite (including adversarial cases) for the generation and agentic layers.
  • A scorer library for grounding, faithfulness, hallucination, and structured-output validity.
  • A reproducible harness wired into CI that blocks releases on quality regressions.
  • Step-level trace capture and scoring for agentic flows, with dashboards leadership can trust.
Required Qualifications
  • Core 4+ years in software / ML engineering, with hands-on work building LLM evaluation or quality tooling.
  • Real understanding of grounding, faithfulness, and hallucination — and how to measure them rigorously.
  • Technical Strong Python and solid engineering practices (reproducibility, CI/CD).
  • Comfort designing evaluation for non-deterministic systems without producing flaky or meaningless metrics.
  • Familiarity with LLM eval frameworks and LLM-as-judge patterns.
  • Nice-to-Have Experience evaluating agentic / multi-step LLM systems.
  • Familiarity with RAG, structured output, and managed LLMs in-VPC.
  • FinTech / financial-services domain or other high-stakes, correctness-critical AI.
  • Background in statistics or measurement / metrics design.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Evaluation & Reliability Engineer
Senior AI Evaluation & Reliability Engineer

Aubergine Solutions Pvt. Ltd. • Ahmedabad District

On-site
INR 3,000,000 - 6,000,000
Great Place To Work certified
AI QA Engineer
AI QA Engineer

Huptech Hr Solutions • Ahmedabad District

On-site
INR 1,200,000 - 2,400,000
AI Engineer
AI Engineer

Andpayments • India

On-site
INR 1,800,000 - 3,000,000
Sr. Agentic AI Engineer
Sr. Agentic AI Engineer

SynapOne • Bengaluru

On-site
INR 900,000 - 1,500,000
Principal Architect AI Data Engineer
Principal Architect AI Data Engineer

EXL • Maharashtra

On-site
INR 3,500,000 - 5,200,000
Agentic AI Engineer - Senior
Agentic AI Engineer - Senior

Interactly.ai • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Applied AI Engineer
Applied AI Engineer

Infer • Karnataka

On-site
INR 1,500,000 - 2,500,000
Senior AI Data Engineer
Senior AI Data Engineer

EXL • Maharashtra

On-site
INR 1,200,000 - 1,800,000
Principal Architect AI Data Engineer
Principal Architect AI Data Engineer

EXL • Gurugram District

On-site
INR 4,000,000 - 9,000,000
QA Engineer
QA Engineer

E2M Solutions • Ahmedabad District

On-site
INR 700,000 - 1,200,000