Senior GenAI Engineer (AI Evaluation Engineer)

FM India

Bengaluru

On-site

INR 1,800,000 - 3,000,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

FM India in Bengaluru seeks an experienced GenAI Evaluation Engineer to lead the design, implementation, and operation of enterprise-grade evaluation and governance frameworks for Generative AI systems in a regulated environment.

You will focus on AI-specific testing, automation, and continuous evaluation pipelines, ensuring quality, safety, and performance of LLMs and related AI workflows in production.

Qualifications

  • Experience evaluating LLMs and AI systems.
  • Knowledge of AI safety, bias, and governance.
  • Familiarity with CI/CD for AI apps.
  • Strong collaboration with product and data teams.

Responsibilities

  • Design AI evaluation and experimentation frameworks.
  • Build automated evaluation systems for outputs.
  • Develop quality benchmarks for safety and compliance.
  • Integrate AI quality gates into CI/CD.

Skills

AI evaluation design
Prompt engineering
CI/CD testing
Monitoring observability
GenAI pipelines
Experimentation frameworks

Education

Bachelor's degree in CS/AI

Tools

Databricks
Azure AI/ML
GCP Vertex
AWS Bedrock

Job description

GenAI Engineer (AI Evaluation Engineer)

Reports To Manager of GenAI Engineering


Job Summary

The Gen AI Evaluation Engineer leads the design, implementation, and operation of enterprise-grade evaluation, quality, and governance frameworks for Generative AI systems in a highly regulated, responsible AI environment. This role ensures the quality, reliability, safety, and performance of LLMs, vision models, RAG pipelines, and agentic workflows deployed in production.Building on strong GenAI engineering foundations, this position focuses on AI-specific testing, experimentation, automation, and continuous evaluation pipelines, integrating quality gates into CI/CD workflows and aligning GenAI solutions with enterprise architecture, compliance, and risk standards. The Gen AI Evaluation Engineer partners closely with product, data science, ML engineering, and platform teams to drive trustworthy, scalable, and production-ready AI systems.


Essential Functions & Responsibilities:


  • GenAI Application Design & Implementation: Design and implement comprehensive AI evaluation and experimentation frameworks for LLMs, vision models, RAG pipelines, and agentic workflows.

  • Build automated evaluation systems to assess model outputs for accuracy, relevance, bias, hallucinations, safety, and regression stability.

  • Develop quality benchmarks and continuous testing pipelines covering content quality, safety, alignment, and enterprise compliance.

  • Establish AI-specific quality gates and acceptance criteria integrated into Agile sprints and CI/CD pipelines.

  • Design, develop, and evaluate data pipelines and RAG workflows using Promptflow, Azure AI Search, ADF Pipelines, Databricks, Spark, and Vector Databases.

  • Validate prompt engineering strategies, prompt consistency, and inference pipelines using GenAI-specific testing tools.

  • Perform prompt-based scenario testing, hallucination detection, and regression validation across model versions.

  • System Support & Operational Excellence: Develop and maintain monitoring capabilities for model drift detection, data quality validation, inference latency, and system reliability.

  • Monitor production GenAI systems using Azure Application Insights, Dynatrace, and AI evaluation platforms.

  • Apply AI risk-focused testing, including bias, fairness, safety, adversarial testing, and red-teaming methodologies.

  • Ensure optimized utilization of infrastructure and compute across diverse AI workloads.

  • Automation & Test Engineering: Build and maintain test automation frameworks (primarily in Python) for AI/ML evaluation and end-to-end workflow validation.

  • Write automation scripts to simulate user behavior and backend interactions.

  • Perform exploratory testing, regression testing, and end-to-end system validation for AI-enabled applications.

  • Track and manage defects using AI evaluation platforms and Agile tools; document test plans, execution results, and evaluation metrics.

  • Mentorship & Team Leadership: Lead design and evaluation reviews; mentor junior engineers and evaluation specialists.

  • Collaborate with product managers to convert business and regulatory requirements into test cases, benchmarks, and test data.

  • Drive innovation by adopting emerging GenAI evaluation techniques and aligning AI initiatives with enterprise goals.

  • Foster a culture of technical excellence, responsible AI, and continuous improvement.


Essential Skills:


  • Design and implementation of AI evaluation frameworks for LLMs, RAG, vision models, and agentic systems

  • Generative AI engineering: prompt engineering, model integration, inference pipelines

  • Data pipelines and RAG workflows (Promptflow, Azure AI Search, ADF Pipelines, Vector DBs)

  • Automated testing and CI/CD integration for AI systems

  • Monitoring and observability (Azure Application Insights, Dynatrace)

  • AI risk, safety, bias, fairness, hallucination detection, and adversarial testing

  • Large-scale experimentation and performance optimization

  • Strong mentorship, technical leadership, and cross-team collaboration


Must Have Skills:


  • Generative AI Azure (AI, Data, and Monitoring services)

  • AI/ML Platform tools (Databricks, Azure AI/ML, GCP Vertex, AWS Bedrock)

  • Python (test automation, evaluation frameworks)

  • Software Engineering & Test Engineering fundamentals


Basic Qualifications:


  • 1-3 years of experience in evaluation or test engineering with a strong focus on AI/ML and Generative AI systems.

  • Hands-on experience evaluating LLMs, RAG pipelines, and AI-powered applications in production environments.

  • Strong understanding of AI/ML evaluation concepts including accuracy, relevance, regression, latency, and stability.

  • Bachelors Degree in Computer Science, Data Science, Artificial Intelligence, or equivalent practical experience.


Preferred Qualifications:


  • Knowledge of AI risks including bias, fairness, safety, hallucinations, and adversarial attacks.

  • Experience in highly regulated or governed AI environments.

  • Strong collaboration skills with data scientists, ML engineers, platform teams, and product managers.

  • Passion for AI quality, responsible AI practices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Generative AI Engineer
Senior Generative AI Engineer

Wissen Infotech • Bengaluru

Hybrid
INR 2,200,000 - 4,000,000
Lead AI Engineer
Lead AI Engineer

Unlimited Innovations • Ernakulam

Hybrid
INR 1,800,000 - 3,000,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Bits In Glass • Hyderabad

On-site
INR 2,000,000 - 3,000,000
Manager AI-ML
Manager AI-ML

Ecolab Global Services • Bengaluru

On-site
INR 2,500,000 - 3,500,000
Gen AI - Engineering Lead
Gen AI - Engineering Lead

ExlService Holdings, Inc. • Dadri

On-site
INR 1,200,000 - 1,800,000
AI Engineer
AI Engineer

Technologies Pvt. Ltd. • Pune District

On-site
INR 1,500,000 - 2,500,000
AI Lead
AI Lead

Tekskills • Gurugram District

On-site
INR 4,200,000 - 8,000,000
Senior AI Security Engineer
Senior AI Security Engineer

Naukri Assist • Pune District, Bengaluru, Delhi

Hybrid
INR 3,000,000 - 5,000,000
Python AI Developer
Python AI Developer

EY • Chennai District, Bengaluru

Hybrid
INR 1,800,000 - 3,200,000
Senior Associate – AI ML Engineer
Senior Associate – AI ML Engineer

Jobtailor • Pune District

On-site
INR 1,200,000 - 2,000,000