Generative AI Quality Engineer - Assistant Vice President

Bot Jobs

Pune District

Hybrid

INR 1,200,000 - 2,400,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Citi invites applications for an AI Quality Engineer to join Retail and Wealth Risk Engineering. You will own AI quality work across agentic AI flow testing, RAG pipelines, AI safety, and test automation at scale.

The role demands deep AI/ML knowledge and breadth in testing disciplines, including evaluation of agent reasoning, memory, tool-use, and multi-agent coordination. Hybrid work in Pune, India, with global impact.

Qualifications

  • Experience in AI quality engineering for autonomous AI systems.
  • Hands-on with agent reasoning, memory, and multi-agent coordination.
  • Familiar with end-to-end RAG pipelines and testing.
  • Proficiency in test automation and CI/CD.

Responsibilities

  • Design and implement end-to-end test strategies for Agentic AI pipelines and multi-agent workflows.
  • Validate agent reasoning, planning, and decision-making chains (ReAct, Chain-of-Thought, Plan-and-Execute).
  • Test tool-use correctness and agent memory systems across sessions.
  • Evaluate context window management and termination conditions in agent loops.
  • Create automated evaluation pipelines integrated with CI/CD.

Skills

Agentic AI Testing
RAG Testing
Test Automation
AI Safety & Security
Multi-agent Frameworks

Tools

Great Expectations
Deequ
LangGraph

Job description

Discover your future at Citi

Working at Citi is far more than just a job. A career with us means joining a team of approximately 219,000 dedicated people from around the globe. At Citi, you'll have the opportunity to grow your career, give back to your community and make a real impact.

Job Overview

We are seeking a highly motivated and experienced AI Quality Engineer to join our Retail and Wealth Risk Engineering team under the Enterprise Risk Technology platform. This role spans the full spectrum of modern AI quality engineering — from Agentic AI flow testing and RAG pipeline validation to AI safety, test automation, and performance & reliability engineering.

You will be the quality pillar for complex autonomous AI systems, ensuring they are safe, accurate, explainable, resilient, and production-ready at scale. This is a high-impact, highly technical role that requires both depth in AI/ML and breadth across testing disciplines.

Job Details

Job Req Id: 26991373

Location(s): Pune, Maharashtra, India, Chennai, Tamil Nadu, India

Job Type: Hybrid

Posted: Sep. 24, 2026

Role Overview

Discover your future at Citi

We are seeking a highly motivated and experienced AI Quality Engineer to join our Retail and Wealth Risk Engineering team under the Enterprise Risk Technology platform. This role spans the full spectrum of modern AI quality engineering — from Agentic AI flow testing and RAG pipeline validation to AI safety, test automation, and performance & reliability engineering.

Agentic AI Testing
  • Design and execute end-to-end test strategies for Agentic AI pipelines, including single-agent and multi-agent workflows.
  • Validate agent reasoning, planning, and decision-making chains (e.g., ReAct, Chain-of-Thought, Plan-and-Execute, Reflexion).
  • Test tool-use correctness — ensuring agents invoke the right tools, with correct parameters, at the right time.
  • Evaluate agent memory systems (short-term, long-term, episodic) for accuracy and context retention across sessions.
  • Validate agent handoff and delegation logic in multi-agent orchestration frameworks (e.g., AutoGen, CrewAI, LangGraph).
  • Test termination conditions, loop detection, and infinite loop prevention in autonomous agent loops.
RAG (Retrieval-Augmented Generation) Testing
  • Design comprehensive test strategies for end-to-end RAG pipelines — covering ingestion, chunking, embedding, retrieval, reranking, and generation stages.
  • Validate retrieval accuracy and relevance — ensuring the correct context chunks are retrieved for a given query.
  • Test embedding model quality and vector similarity thresholds across different document corpora.
  • Evaluate faithfulness, groundedness, and answer relevance of generated responses using frameworks like RAGAS, TruLens, DeepEval.
  • Test chunking strategies (fixed, semantic, hierarchical) for their impact on retrieval quality.
  • Validate context window management — ensuring retrieved context does not exceed token limits or degrade generation quality.
  • Conduct end-to-end regression testing when the underlying knowledge base, embedding model, or LLM changes.
  • Test multi-turn conversational RAG for context coherence and citation accuracy across turns.
Test Automation
  • Build and maintain automated test harnesses for Agentic and RAG systems, including agent trajectory replay, tool mock injection, and prompt simulation.
  • Develop automated evaluation pipelines integrated into CI/CD workflows for continuous model and agent validation.
  • Create data validation and data quality frameworks using Great Expectations, Deequ, or custom tooling for training, retrieval, and inference data.
  • Build prompt regression suites to detect behavioral drift across LLM versions or prompt changes.
  • Implement determinism and reproducibility tests for stochastic LLM-based decisions.
  • Automate vector database validation — index integrity, embedding drift, and retrieval consistency checks.
AI Safety & Security Testing
  • Conduct red-teaming and adversarial testing to uncover jailbreaks, prompt injection vulnerabilities, and goal misalignment in LLM-based systems.
  • Test output guardrails and content filters for unsafe, biased, toxic, or out-of-scope model behavior.
  • Validate privilege escalation controls — ensuring agents do not exceed permitted actions or access unauthorized resources.
  • Perform data poisoning and backdoor attack simulations to assess model robustness.
  • Evaluate models for bias, fairness, and discrimination using frameworks such as AI Fairness 360 and Aequitas.
  • Test PII leakage and data privacy controls in RAG and agent pipelines in accordance with GDPR,
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Generative AI Quality Engineer - Assistant Vice President
Generative AI Quality Engineer - Assistant Vice President

Citigroup Inc. • Pune District

On-site
INR 1,800,000 - 3,200,000
AI Quality Engineer, Safety and RAG - Assistant Vice President
AI Quality Engineer, Safety and RAG - Assistant Vice President

Citi • Pune District

On-site
INR 900,000 - 1,400,000
AI Engineer, Agentic Systems (Quality Engineering)
AI Engineer, Agentic Systems (Quality Engineering)

Bot Jobs • Bengaluru

On-site
INR 1,800,000 - 2,400,000
AI Quality Engineer, Safety and RAG – Assistant Vice President
AI Quality Engineer, Safety and RAG – Assistant Vice President

Growth For Impact • Chennai District

On-site
INR 900,000 - 1,300,000
Generative AI Quality Engineer - Assistant Vice President
Generative AI Quality Engineer - Assistant Vice President

Citi • Maharashtra

On-site
INR 250,000 - 500,000
AI QE Engineer - Manager
AI QE Engineer - Manager

PwC • Hyderabad, Chennai District, Bengaluru

On-site
INR 1,500,000 - 2,100,000
AI Quality Automation Engineer – Assistant Vice President
AI Quality Automation Engineer – Assistant Vice President

Citigroup Inc. • Pune District

On-site
INR 3,000,000 - 6,000,000
AI Quality Automation Engineer – Assistant Vice President
AI Quality Automation Engineer – Assistant Vice President

Citi • Maharashtra

On-site
INR 2,500,000 - 4,500,000
Quality Assurance (QA) Engineer
Quality Assurance (QA) Engineer

Aganitha Cognitive Solutions • Chennai District

On-site
INR 900,000 - 1,500,000
Principal Software Engineer
Principal Software Engineer

Cadence • Bengaluru

On-site
INR 4,500,000 - 7,500,000