Generative AI Quality Engineer - Assistant Vice President

Citigroup Inc.

Pune District

On-site

INR 1,800,000 - 3,200,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Citigroup Inc. is seeking an experienced AI Quality Engineer to lead testing of Agentic AI and RAG pipelines. You will own end-to-end test strategies, validate reasoning and tool use, and drive automation across safety and reliability domains.

The role emphasizes production-ready AI systems, performance optimization, and collaboration with MLOps and security teams in a fast-paced environment.

Qualifications

  • 8+ years in software or AI/ML quality engineering.
  • Experience building automated test frameworks for non-deterministic AI systems.
  • Strong background in performance testing and AI safety assessments.
  • Experience with RAG systems and agentic AI testing is highly preferred.
  • Familiarity with CI/CD workflows for model and agent validation.

Responsibilities

  • Design and execute end-to-end test strategies for Agentic AI pipelines, including multi-agent workflows.
  • Validate agent reasoning, planning, and decision-making chains (e.g., ReAct, Chain-of-Thought).
  • Test tool-use correctness and verify memory and handoff across agents.
  • Develop automated test harnesses and CI/CD pipelines for AI systems.
  • Conduct red-teaming, security testing, and bias/fairness assessments for AI components.
  • Define load, stress, and reliability tests; establish SLOs/SLAs for AI services.

Skills

Agentic AI Testing
RAG Testing
Test Automation
AI Safety Testing
Performance Testing
Reliability Engineering
CI/CD

Education

Bachelor's degree in Computer Science or related field
Master's degree is a plus

Tools

Great Expectations
Deequ
TruLens
DeepEval
RAGAS
AutoGen
CrewAI
LangGraph

Job description

We are seeking a highly motivated and experienced AI Quality Engineer to join our Retail and Wealth Risk Engineering team under the Enterprise Risk Technology platform. This role spans the full spectrum of modern AI quality engineering — from Agentic AI flow testing and RAG pipeline validation to AI safety, test automation, and performance & reliability engineering.

You will be the quality pillar for complex autonomous AI systems, ensuring they are safe, accurate, explainable, resilient, and production-ready at scale. This is a high-impact, highly technical role that requires both depth in AI/ML and breadth across testing disciplines.

Responsibilities
AgenticAI Testing
  • Design and executeend-to-end test strategies for Agentic AI pipelines, including single-agent and multi-agent workflows.
  • Validateagent reasoning, planning, and decision-making chains(e.g., ReAct, Chain-of-Thought, Plan-and-Execute, Reflexion).
  • Testtool-use correctness— ensuring agents invoke the right tools, with correct parameters, at the right time.
  • Evaluateagent memory systems(short-term, long-term, episodic) for accuracy and context retention across sessions.
  • Validateagent handoff and delegation logicin multi-agent orchestration frameworks (e.g., AutoGen, CrewAI, LangGraph).
  • Testtermination conditions, loop detection, andinfinite loop preventionin autonomous agent loops.
RAG (Retrieval-Augmented Generation) Testing
  • Design comprehensive test strategies forend-to-end RAG pipelines— covering ingestion, chunking, embedding, retrieval, reranking, and generation stages.
  • Validateretrieval accuracy and relevance— ensuring the correct context chunks are retrieved for a given query.
  • Testembedding model qualityand vector similarity thresholds across different document corpora.
  • Evaluatefaithfulness, groundedness, and answer relevanceof generated responses using frameworks likeRAGAS, TruLens, DeepEval.
  • Testchunking strategies(fixed, semantic, hierarchical) for their impact on retrieval quality.
  • Validatecontext window management— ensuring retrieved context does not exceed token limits or degrade generation quality.
  • Conductend-to-end regression testingwhen the underlying knowledge base, embedding model, or LLM changes.
  • Testmulti-turn conversational RAGfor context coherence and citation accuracy across turns.
Test Automation
  • Build and maintainautomated test harnessesfor Agentic and RAG systems, including agent trajectory replay, tool mock injection, and prompt simulation.
  • Developautomated evaluation pipelinesintegrated into CI/CD workflows for continuous model and agent validation.
  • Createdata validation and data quality frameworks(using Great Expectations, Deequ, or custom tooling) for training, retrieval, and inference data.
  • Buildprompt regression suitesto detect behavioral drift across LLM versions or prompt changes.
  • Implementdeterminism and reproducibility testsfor stochastic LLM-based decisions.
  • Automatevector database validation— index integrity, embedding drift, and retrieval consistency checks.
AI Safety & Security Testing
  • Conductred-teaming and adversarial testingto uncover jailbreaks, prompt injection vulnerabilities, and goal misalignment in LLM-based systems.
  • Testoutput guardrails and content filtersfor unsafe, biased, toxic, or out-of-scope model behavior.
  • Validateprivilege escalation controls— ensuring agents do not exceed permitted actions or access unauthorized resources.
  • Performdata poisoning and backdoor attack simulationsto assess model robustness.
  • Evaluate models forbias, fairness, and discriminationusing frameworks such as AI Fairness 360 and Aequitas.
  • TestPII leakage and data privacy controlsin RAG and agent pipelines in accordance with GDPR, CCPA, and internal data governance policies.
  • Conduct security testing aligned with theOWASP Top 10 for LLM Applications, including:
    • Prompt Injection (Direct & Indirect)
    • Insecure Output Handling
    • Training Data Poisoning
    • Insecure Plugin / Tool Design
    • Sensitive Information Disclosure
  • Validateconstitutional AI constraints, RLHF-aligned behavior boundaries, and system prompt integrity.
  • Collaborate with cybersecurity teams onAI-specific threat modelingand vulnerability management.
  • Maintainsafety testing playbooksand document red-team findings with severity ratings and remediation recommendations.
Performance & Reliability Testing
  • Define and executeload, stress, soak, and spike testingfor AI-powered APIs, inference endpoints, and agent orchestration services.
  • Measure and optimizeend-to-end latencyacross RAG and agentic pipelines — from query to final response.
  • BenchmarkLLM inference throughput(tokens/second) and identify bottlenecks across model serving infrastructure.
  • Testauto-scaling behaviorof AI services under variable load conditions.
  • Validatecircuit breaker, retry, and fallback mechanismsin agentic and RAG systems for graceful degradation.
  • Testvector database performance— query latency, index build time, and retrieval accuracy under high concurrency.
  • Conductcost efficiency analysis— measuring token consumption, API call costs, and infrastructure spend per agent task.
  • EstablishSLOs (Service Level Objectives)andSLAsfor AI system availability, latency percentiles (P50, P95, P99), and error rates.
  • Collaborate with MLOps teams to set upobservability dashboards, monitoring alerts, and automated anomaly detection for production AI systems.
  • Performchaos engineering experimentsto validate agent and RAG system resilience under infrastructure failures.
Domain Knowledge
  • Deep understanding ofRAG architecture patterns— naive RAG, advanced RAG, modular RAG.
  • Solid grasp ofagent design patterns: ReAct, Plan-and-Execute, Reflexion, MRKL, Mixture-of-Agents.
  • Familiarity withAI safety and alignmentprinciples (RLHF, Constitutional AI, guardrail layers).
  • Knowledge oftoken economics, context management, and LLM cost optimization.
  • Proficiency inperformance engineeringmethodologies for distributed AI systems.
Preferred Qualifications
  • Experience withMCP (Model Context Protocol)or similar agentic communication standards.
  • Exposure tomulti-modal agent testing(agents handling text, images, code, documents).
  • Experience inregulated industries(banking, finance, healthcare) with strict compliance requirements.
  • Familiarity withchaos engineeringtools (Chaos Monkey, Gremlin, LitmusChaos).
Education
  • Bachelor’s degree in Computer Science, Engineering, or a related field.
  • Master’s degree is a plus.
Experience
  • 8+ years of experience in software or AI/ML quality engineering.
  • 3+ years of hands-on experience with RAG systems, or Agentic AI.
  • Proven experience building automated test frameworks for non-deterministic AI systems.
  • Strong background in performance testing and AI safety/security assessments.
Job Family Group

Technology

Job Family

Technology Quality

Time Type

Full time

Most Relevant Skills

Please see the requirements listed above.

Other Relevant Skills

For complementary skills, please see above.

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Quality Engineer, Safety and RAG - Assistant Vice President
AI Quality Engineer, Safety and RAG - Assistant Vice President

Citi • Pune District

On-site
INR 900,000 - 1,400,000
Generative AI Quality Engineer - Assistant Vice President
Generative AI Quality Engineer - Assistant Vice President

Citi • Maharashtra

On-site
INR 250,000 - 500,000
AI Quality Engineer, Safety and RAG - Assistant Vice President
AI Quality Engineer, Safety and RAG - Assistant Vice President

Citigroup Inc. • Chennai District

On-site
INR 3,000,000 - 6,000,000
Generative AI Quality Engineer - Assistant Vice President
Generative AI Quality Engineer - Assistant Vice President

Bot Jobs • Pune District

On-site
INR 1,200,000 - 2,400,000
Generative AI Quality Engineer - Assistant Vice President
Generative AI Quality Engineer - Assistant Vice President

Citibank (Switzerland) AG • Pune District

On-site
INR 1,800,000 - 3,000,000
AI Quality Automation Engineer – Assistant Vice President
AI Quality Automation Engineer – Assistant Vice President

Citigroup Inc. • Pune District

On-site
INR 3,000,000 - 6,000,000
AI Quality Automation Engineer – Assistant Vice President
AI Quality Automation Engineer – Assistant Vice President

Citi • Maharashtra

On-site
INR 2,500,000 - 4,500,000
AI Quality Automation Engineer – Assistant Vice President
AI Quality Automation Engineer – Assistant Vice President

Citibank (Switzerland) AG • Pune District

On-site
Confidential
Senior QA Automation Engineer (Python, Selenium & AI Testing) - Assistant Vice President
Senior QA Automation Engineer (Python, Selenium & AI Testing) - Assistant Vice President

Citi • Maharashtra

On-site
INR 1,200,000 - 1,800,000
AI Engineer - Assistant Vice President
AI Engineer - Assistant Vice President

Citi • Maharashtra

On-site
INR 1,500,000 - 2,200,000