QA Lead- AI Evaluation & Quality

National e-Governance Division

New Delhi

On-site

INR 1,500,000 - 1,900,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

National e-Governance Division seeks a seasoned AI quality engineer to lead evaluation strategy for government‑grade ML/AI systems. You will define formal gates, build reusable evaluation harnesses, and ensure adherence to responsible AI guidelines across ministry deployments.

The role requires 5+ years in ML QA, hands-on experience with LLM/RAG evaluation, and embedding tests in CI/CD. Prior government experience is a strong plus; a structured, governance minded approach is essential.

Qualifications

  • 5+ years in ML or software quality assurance or applied AI.
  • Minimum 3 years in AI/ML model evaluation or ML QA.
  • Hands-on experience building evaluation harnesses and automated suites.
  • Experience evaluating LLM/RAG systems for hallucination and retrieval quality.
  • Experience embedding evaluations into CI/CD pipelines.
  • Experience in government/regulatory sectors is a strong plus.
  • Desirable: directing/mentoring QA/evaluation standards.

Responsibilities

  • Own the programme's AI evaluation methodology and gate in delivery lifecycle.
  • Choose evaluation approaches per use case and document rationale.
  • Ensure Responsible AI evaluation aligns with MeitY and IndiaAI frameworks.
  • Design automated evaluation harnesses, datasets and regression suites.
  • Wire evaluations into CI/CD as release gates and version data/prompts.
  • Publish reusable evaluation assets to AIKosh/OpenForge for national reuse.
  • Assess models with metrics and report with statistical rigor.

Skills

ML QA
Model evaluation
CI/CD integration
Evaluation governance
Python
SQL
Regulated sector experience

Education

B.Tech./B.E. or M.Tech./M.S./M.Sc. in CS/DS/AI
Advanced degree desirable (MTech/MS/PhD)
Certifications: ISTQB Foundation, CT-AI
Cloud/MLOps certification desirable

Tools

CI/CD pipelines
Python
SQL

Job description

Educational Qualification
  • B.Tech./B.E. or M.Tech./M.S./M.Sc. in Computer Science, Data Science, AI/ML, or a related quantitative discipline ( Must have )
  • Advanced degree (M.Tech./M.S./Ph.D.) or published/applied work in ML evaluation, NLP/LLM evaluation, or Responsible AI is highly desirable
  • Certifications (Desirable): ISTQB Foundation and ISTQB Certified Tester AI Testing (CT-AI); DeepLearning.AI or equivalent ML foundations; a cloud (AWS/Azure/GCP) or MLOps certification
  • Non-traditional backgrounds with demonstrable model-evaluation or ML-quality work, open-source evaluation or data-quality contributions, or competitive benchmark/Kaggle achievements will be considered in lieu of formal qualification

Experience
  • 5+ years in ML or software quality assurance, test automation, data or ML engineering, or applied AI; minimum 3 years specifically in AI/ML model evaluation, LLM/Gen-AI evaluation, or ML quality assurance
  • Demonstrable hands‑on experience building evaluation harnesses, benchmark/golden datasets and automated evaluation suites for ML or LLM systems
  • Demonstrable experience evaluating LLM/RAG systems (hallucination, groundedness, retrieval quality) and/or agentic systems (tool‑use correctness, trajectory, task completion)
  • Prior experience embedding evaluations into CI/CD pipelines and versioning datasets, prompts and experiments
  • Prior experience delivering AI quality or evaluation work in government, BFSI, healthcare, or another regulated sector is a strong plus
  • Prior experience directing, mentoring, or setting QA/evaluation standards across multiple teams is desirable, given the programme‑lead nature of the role

Key Responsibilities
Evaluation Strategy, Standards & Quality Gates
  • Own the programmes AI evaluation methodology and make evaluation a formal gate in the delivery lifecycle defining, publishing and enforcing common metrics, acceptance thresholds and quality gates
  • Select the appropriate evaluation approach per use case — offline benchmarks, human evaluation, LLM‑as‑judge, or online/production evaluation — and document the rationale
  • Ensure Responsible AI evaluation is built into the standard delivery flow rather than added afterwards, aligned to the MeitY Responsible AI advisory and the IndiaAI Safe & Trusted AI framework
Evaluation Infrastructure & Reusable Harnesses
  • Design and build automated evaluation harnesses, benchmark and golden datasets, and regression suites that pods reuse across ministry deployments — mirroring the programme’s build‑once, reuse‑many model
  • Wire evaluations into CI/CD as release gates; version datasets, prompts and evaluation configurations so results are reproducible and audit‑ready
  • Publish reusable evaluation sets, rubrics, harnesses and templates to AIKosh under standard metadata (and to OpenForge where code is shared) for national reuse
Model & Output Quality Evaluation
  • Evaluate models with task‑appropriate quantitative metrics — precision/recall/F1, ROC‑AUC/PR‑AUC, calibration, MAE/RMSE and task‑specific measures — reported with statistical rigour, including uncertainty, adequate sample sizes and vigilance against metric gaming
  • Evaluate AI outputs against defined quality dimensions — accuracy, groundedness, reliability, consistency, relevance and adherence to business and policy requirements — rather than pass/fail alone
Gen‑AI, LLM & RAG Evaluation
  • Assess LLM and RAG systems for hallucination, factual consistency, groundedness and citation faithfulness, relevance, prompt robustness, input validation and non‑deterministic behaviour
  • Measure retrieval quality and generation quality separately in RAG pipelines using RAGAS, DeepEval, TruLens or LangSmith
  • Use LLM‑as‑judge where appropriate, calibrated against human labels, controlling for known judge biases (position, verbosity, self‑preference), and defaulting to human evaluation where automated judgement is unreliable
Agentic AI Evaluation
  • Evaluate agents on their trajectories, not only final answers — tool selection and tool‑call correctness, multi‑step reasoning and workflow completion, agent handoffs, failure recovery, memory behaviour and end‑to‑end task‑completion accuracy
  • Validate human‑in‑the‑loop controls — review, escalation, override, feedback and exception handling — wherever AI decisions require human supervision
Safety, Adversarial & Responsible AI Testing
  • Under the technical guidance of the AI Safety Researcher, operationalise adversarial and red‑team findings — jailbreaks, prompt injection, indirect prompt injection, unsafe behaviour and sensitive‑data (PII) leakage — into standard, reusable test suites the pods apply, and verify guardrail effectiveness (Llama Guard, NeMo Guardrails, Guardrails AI or equivalent)
  • Evaluate fairness by measuring performance across demographic, linguistic, geographic and other relevant cohorts using appropriate fairness metrics (Fairlearn, AI Fairness 360)
  • Author the QA‑side safety and quality evidence — model cards, dataset sheets, evaluation logs and Responsible AI records — for approval
Data & Dataset Quality Validation
  • Validate training, evaluation and inference data for missing values, duplicates, schema conformance, class imbalance, data leakage, distribution/covariate shift, feature consistency and train–serve skew, using tools such as Great Expectations, Evidently or Deepchecks
  • Curate high‑quality evaluation datasets that reflect real citizen usage, edge cases and Indian‑language and regional diversity
Production Monitoring & Quality Operations
  • Monitor live AI applications for quality and retrieval degradation, drift (data, concept and model), accuracy and latency regressions, token/cost behaviour, failures and safety/compliance violations
  • Define alerts, thresholds and dashboards for model‑quality, latency, error‑rate, drift and compliance signals; support quality‑incident response and feed production findings back into the evaluation sets
Human Evaluation Operations
  • Design evaluation rubrics and run structured human evaluation and annotation; manage annotator quality and inter‑annotator agreement; build preference/label datasets that ground and validate the automated metrics
Cross‑Pod Leadership, Mentoring & Governance
  • Provide technical direction to the AI QA Engineers in the team; set shared standards and harnesses, keep pod‑level practice consistent and current, and drive adoption of sound, current evaluation techniques while screening out hype
  • Turn quality evidence into decision‑ready go/no‑go recommendations for deployments, and represent AI quality in programme governance and monthly reviews
  • Contribute AI quality and evaluation requirements to NeGD RFQs and RDRs, and review agency deliverables for evaluation and quality posture

Technical Competencies
  • Programming & Data: Python (advanced) and SQL; pandas, NumPy, scikit‑learn; pytest for test authoring
  • Evaluation & Measurement: offline benchmarks, human evaluation, LLM‑as‑judge and online/production evaluation; statistical rigour — significance testing, bootstrap confidence intervals, calibration (ECE) and inter‑annotator agreement (Cohen’s / Fleiss’ kappa)
  • LLM, RAG & Agentic Evaluation: RAGAS, DeepEval, TruLens, LangSmith or Promptfoo; separate retrieval‑ and generation‑quality measurement; trajectory and tool‑use evaluation for agents
  • ML Metrics: precision/recall/F1, ROC‑AUC and PR‑AUC, calibration, MAE/RMSE and task‑specific metrics
  • Data & Dataset Quality: Great Expectations, Evidently, Deepchecks or TFDV; detection of leakage, drift and covariate shift, class imbalance, schema violations and train–serve skew
  • Safety & Responsible AI Testing: guardrail testing (Llama Guard, NeMo Guardrails, Guardrails AI), PII‑leakage and unsafe‑output testing; fairness metrics (Fairlearn, AI Fairness 360); model cards and dataset sheets — applied under the AI Safety Researcher’s guidance
  • MLOps & Observability: MLflow or Weights & Biases for experiment and dataset versioning; Prometheus, Grafana, OpenTelemetry; LLM observability (Langfuse, Arize Phoenix or equivalent); CI/CD (Jenkins, GitHub Actions)
  • Test Automation (supporting): pytest as the primary automation stack; Postman or REST‑assured for API testing; Selenium or Cypress where UI validation is relevant
  • Security (AI‑aware): PII/sensitive‑data exposure, access controls around AI services, model/API misuse and data leakage, in addition to conventional application‑security fundamentals (authentication/authorisation, encryption, secure APIs)
  • Cloud & Platforms: AWS, Azure, GCP and IndiaAI Compute evaluation and testing environments; containerised evaluation (Docker, Kubernetes)
  • Government Frameworks & Standards: IndiaAI Safe & Trusted AI framework, MeitY Responsible AI advisory, MeitY security guidelines, CERT‑In directions, DPDPA 2023
  • Communication & Leadership: ability to author clear, decision‑ready evaluation and quality reports for engineering, architecture and executive audiences; to set standards and mentor pod QA engineers; and to represent AI quality in cross‑ministerial reviews

Desirable Evidence at Screening
  • A built evaluation harness, automated evaluation suite, or benchmark/golden dataset for an ML or LLM system, with a verifiable repository, notebook or demonstrable artefact
  • Documented model or LLM/RAG evaluation with quantitative results against a production or production‑like system
  • Open‑source contribution to an ML/LLM evaluation, data‑quality, or AI‑observability project with verifiable commit history
  • Prior AI evaluation or quality‑assurance engagement for a government, BFSI, healthcare or regulated‑sector client, verifiable via reference or letter of confirmation
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Safety Research & Testing Professional
AI Safety Research & Testing Professional

Data Security Council of India • Delhi

On-site
INR 1,800,000 - 2,800,000
AI Quality Assurance Engineer
AI Quality Assurance Engineer

Thirdeye Data Inc. • Bengaluru

Hybrid
INR 1,200,000 - 2,000,000
AI Safety Research & Testing Professional
AI Safety Research & Testing Professional

Data Security Council of India • Dadri

On-site
INR 900,000 - 1,500,000
QA Engineer
QA Engineer

YO IT Consulting • Mumbai

Hybrid
INR 800,000 - 1,400,000
Vice President AI/ ML
Vice President AI/ ML

National e-Governance Division • New Delhi

On-site
INR 900,000 - 1,300,000
Vice President – AI/ML
Vice President – AI/ML

National e-Governance Division • New Delhi

On-site
INR 4,000,000 - 8,000,000
Quality Assurance Engineer
Quality Assurance Engineer

Valiance Solutions • Dadri

On-site
INR 800,000 - 1,500,000
Senior AI Evaluation & Reliability Engineer
Senior AI Evaluation & Reliability Engineer

Aubergine Solutions Pvt. Ltd. • Ahmedabad District

On-site
INR 3,000,000 - 6,000,000
Great Place To Work certified
AI Trust and Governance Architect
AI Trust and Governance Architect

Infosys • Bengaluru

On-site
INR 1,600,000 - 2,400,000
AI QA Engineer
AI QA Engineer

Huptech Hr Solutions • Ahmedabad District

On-site
INR 1,200,000 - 2,400,000