Generative AI Quality Engineer - Assistant Vice President

Citi

Maharashtra

On-site

INR 250,000 - 500,000

Full time

3 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Citi is seeking a highly skilled AI Quality Engineer to join the Retail and Wealth Risk Engineering team. You will own end-to-end QA for Agentic AI and Retrieval-Augmented Generation pipelines, ensuring safety, accuracy, and reliability at scale.

The role requires deep AI/ML testing expertise and broad knowledge across automation, security, and performance disciplines. Responsibilities span designing test strategies, validating reasoning and tool use, and implementing automated evaluation within

Qualifications

  • 8+ years of experience in software or AI/ML quality engineering.
  • Experience building automated test frameworks for non-deterministic AI systems.
  • Proven ability to design end-to-end QA for Agentic AI and RAG pipelines.

Responsibilities

  • Design and execute end-to-end test strategies for Agentic AI pipelines.
  • Validate agent reasoning, planning, and decision-making chains.
  • Test tool-use correctness and memory systems across sessions.
  • Evaluate multi-agent handoff in orchestration frameworks.
  • Design test strategies for RAG pipelines including retrieval, embedding and generation stages.
  • Develop automated evaluation pipelines for CI/CD workflows.
  • Perform data poisoning and security testing to assess robustness.
  • Define SLOs/SLAs and collaborate with MLOps for observability.

Skills

AI/ML quality engineering
Test automation
Performance testing
Safety testing

Education

Bachelor’s degree in CS/Engineering
Master’s degree (plus)

Tools

Great Expectations
Deequ

Job description

  • Training Data Poisoning

We are seeking a highly motivated and experienced AI Quality Engineer to join our Retail and Wealth Risk Engineering team under the Enterprise Risk Technology platform. This role spans the full spectrum of modern AI quality engineering — from Agentic AI flow testing and RAG pipeline validation to AI safety, test automation, and performance & reliability engineering.

You will be the quality pillar for complex autonomous AI systems, ensuring they are safe, accurate, explainable, resilient, and production-ready at scale. This is a high-impact, highly technical role that requires both depth in AI/ML and breadth across testing disciplines.

Responsibilities
Agentic AI Testing
  • Design and execute end-to-end test strategies for Agentic AI pipelines, including single-agent and multi-agent workflows.
  • Validate agent reasoning, planning, and decision-making chains (e.g., ReAct, Chain-of-Thought, Plan-and-Execute, Reflexion).
  • Test tool-use correctness — ensuring agents invoke the right tools, with correct parameters, at the right time.
  • Evaluate agent memory systems (short-term, long-term, episodic) for accuracy and context retention across sessions.
  • Validate agent handoff and delegation logic in multi-agent orchestration frameworks (e.g., AutoGen, CrewAI, LangGraph).
  • Test termination conditions, loop detection, and infinite loop prevention in autonomous agent loops.
RAG (Retrieval-Augmented Generation) Testing
  • Design comprehensive test strategies for end-to-end RAG pipelines — covering ingestion, chunking, embedding, retrieval, reranking, and generation stages.
  • Validate retrieval accuracy and relevance — ensuring the correct context chunks are retrieved for a given query.
  • Test embedding model quality and vector similarity thresholds across different document corpora.
  • Evaluate faithfulness, groundedness, and answer relevance of generated responses using frameworks like RAGAS, TruLens, DeepEval.
  • Test chunking strategies (fixed, semantic, hierarchical) for their impact on retrieval quality.
  • Validate context window management — ensuring retrieved context does not exceed token limits or degrade generation quality.
  • Conduct end-to-end regression testing when the underlying knowledge base, embedding model, or LLM changes.
  • Test multi-turn conversational RAG for context coherence and citation accuracy across turns.
Test Automation
  • Build and maintain automated test harnesses for Agentic and RAG systems, including agent trajectory replay, tool mock injection, and prompt simulation.
  • Develop automated evaluation pipelines integrated into CI/CD workflows for continuous model and agent validation.
  • Create data validation and data quality frameworks (using Great Expectations, Deequ, or custom tooling) for training, retrieval, and inference data.
  • Build prompt regression suites to detect behavioral drift across LLM versions or prompt changes.
  • Implement determinism and reproducibility tests for stochastic LLM-based decisions.
  • Automate vector database validation — index integrity, embedding drift, and retrieval consistency checks.
AI Safety & Security Testing
  • Conduct red-teaming and adversarial testing to uncover jailbreaks, prompt injection vulnerabilities, and goal misalignment in LLM-based systems.
  • Test output guardrails and content filters for unsafe, biased, toxic, or out-of-scope model behavior.
  • Validate privilege escalation controls — ensuring agents do not exceed permitted actions or access unauthorized resources.
  • Perform data poisoning and backdoor attack simulations to assess model robustness.
  • Evaluate models for bias, fairness, and discrimination using frameworks such as AI Fairness 360 and Aequitas.
  • Test PII leakage and data privacy controls in RAG and agent pipelines in accordance with GDPR, CCPA, and internal data governance policies.
  • Conduct security testing aligned with the OWASP Top 10 for LLM Applications, including:
    • Prompt Injection (Direct & Indirect)
    • Insecure Output Handling
    • Training Data Poisoning
    • Insecure Plugin / Tool Design
    • Sensitive Information Disclosure
  • Validate constitutional AI constraints, RLHF-aligned behavior boundaries, and system prompt integrity.
  • Collaborate with cybersecurity teams on AI-specific threat modeling and vulnerability management.
  • Maintain safety testing playbooks and document red-team findings with severity ratings and remediation recommendations.
Performance & Reliability Testing
  • Define and execute load, stress, soak, and spike testing for AI-powered APIs, inference endpoints, and agent orchestration services.
  • Measure and optimize end-to-end latency across RAG and agentic pipelines — from query to final response.
  • Benchmark LLM inference throughput (tokens/second) and identify bottlenecks across model serving infrastructure.
  • Test auto-scaling behavior of AI services under variable load conditions.
  • Validate circuit breaker, retry, and fallback mechanisms in agentic and RAG systems for graceful degradation.
  • Test vector database performance — query latency, index build time, and retrieval accuracy under high concurrency.
  • Conduct cost efficiency analysis — measuring token consumption, API call costs, and infrastructure spend per agent task.
  • Establish SLOs (Service Level Objectives) and SLAs for AI system availability, latency percentiles (P50, P95, P99), and error rates.
  • Collaborate with MLOps teams to set up observability dashboards, monitoring alerts, and automated anomaly detection for production AI systems.
  • Perform chaos engineering experiments to validate agent and RAG system resilience under infrastructure failures.
Domain Knowledge
  • Deep understanding of RAG architecture patterns — naive RAG, advanced RAG, modular RAG.
  • Solid grasp of agent design patterns: ReAct, Plan-and-Execute, Reflexion, MRKL, Mixture-of-Agents.
  • Familiarity with AI safety and alignment principles (RLHF, Constitutional AI, guardrail layers).
  • Knowledge of token economics, context management, and LLM cost optimization.
  • Proficiency in performance engineering methodologies for distributed AI systems.
Preferred Qualifications
  • Experience with MCP (Model Context Protocol) or similar agentic communication standards.
  • Exposure to multi-modal agent testing (agents handling text, images, code, documents).
  • Experience in regulated industries (banking, finance, healthcare) with strict compliance requirements.
  • Familiarity with chaos engineering tools (Chaos Monkey, Gremlin, LitmusChaos).
Education
  • Bachelor’s degree in Computer Science, Engineering, or a related field.
  • Master’s degree is a plus.
Experience
  • 8+ years of experience in software or AI/ML quality engineering.
  • 3+ years of hands-on experience with RAG systems, or Agentic AI.
  • Proven experience building automated test frameworks for non-deterministic AI systems.
  • Strong background in performance testing and AI safety/security assessments.
Job Family Group

Technology

Job Family

Technology Quality

Time Type

Full time

Most Relevant Skills

Please see the requirements listed above.

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Quality Engineer, Safety and RAG - Assistant Vice President
AI Quality Engineer, Safety and RAG - Assistant Vice President

Citigroup Inc. • Chennai District

On-site
INR 3,000,000 - 6,000,000
AI Quality Engineer, Safety and RAG - Assistant Vice President
AI Quality Engineer, Safety and RAG - Assistant Vice President

Citi • Pune District

On-site
INR 900,000 - 1,400,000
Generative AI Quality Engineer - Assistant Vice President
Generative AI Quality Engineer - Assistant Vice President

Citigroup Inc. • Pune District

On-site
INR 1,800,000 - 3,200,000
Generative AI Quality Engineer - Assistant Vice President
Generative AI Quality Engineer - Assistant Vice President

Citibank (Switzerland) AG • Pune District

On-site
INR 1,800,000 - 3,000,000
Senior Agentic AI Engineer - Vice President
Senior Agentic AI Engineer - Vice President

Citi • Maharashtra

On-site
INR 400,000 - 700,000
AI Quality Automation Engineer – Assistant Vice President
AI Quality Automation Engineer – Assistant Vice President

Citibank (Switzerland) AG • Pune District

On-site
Confidential
Senior AI Engineer - Vice President
Senior AI Engineer - Vice President

Citigroup Inc. • Pune District

On-site
INR 3,000,000 - 6,000,000
Senior LLM and Agentic AI Engineer - Assistant Vice President
Senior LLM and Agentic AI Engineer - Assistant Vice President

Citi • Chennai District

On-site
INR 4,000,000 - 7,000,000
Senior QA Automation Engineer (Python, Selenium & AI Testing) - Assistant Vice President
Senior QA Automation Engineer (Python, Selenium & AI Testing) - Assistant Vice President

Citi • Pune District

On-site
INR 1,800,000 - 3,000,000
AI Engineer - Assistant Vice President
AI Engineer - Assistant Vice President

Citi • Pune District

On-site
INR 2,500,000 - 5,000,000