Director, Product Management AI Quality & Evaluation

FactSet Research Systems Inc.

Hyderabad

On-site

INR 4,000,000 - 7,000,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

FactSet Research Systems Inc. is seeking a seasoned Product Manager to lead AI/evaluation platform strategy. You will own the evaluation framework, define golden pathways, and partner with engineering to scale enterprise-grade quality across the product portfolio.

You will drive adoption across product teams, define metrics and taxonomies, manage human review operations, and ensure a reproducible path from production to regression tests, contributing to safe and high-quality AI at scale.

Qualifications

  • 10+ years in product management with AI/ML or data-intensive products
  • Experience defining evaluation methodology for ML/AI systems
  • Ability to drive platform adoption across multiple product teams
  • Strong communication and change-management skills
  • Comfort influencing senior stakeholders across time zones

Responsibilities

  • Define and own the AI evaluation platform and its roadmap
  • Set golden pathways for experimentation and release processes
  • Lead human review operations and calibration datasets
  • Drive enterprise adoption and enablement across product teams
  • Close the loop from production failures to regression tests

Skills

Product management
AI/ML
Cross-functional leadership
Stakeholder influence
Written communication

Tools

LangSmith
LangFuse
Braintrust
Weights & Biases

Job description

FactSet creates flexible, open data and software solutions for over 200,000 investment professionals worldwide, providing instant access to financial data and analytics that investors use to make crucial decisions.At FactSet, our values are the foundation of everything we do. They express how we act and operate, serve as a compass in our decision-making, and play a big role in how we treat each other, our clients, and our communities. We believe that the best ideas can come from anyone, anywhere, at any time, and that curiosity is the key to anticipating our clients’ needs and exceeding their expectations.We ship AI into workflows where being confidently wrong is expensive. A banker builds a pitchbook from generated tombstones and charts. A portfolio manager acts on AI-generated attribution commentary, and an analyst asks a conversational assistant a question and gets back an answer synthesized across filings, transcripts, estimates, news, and their own firm's internal data. In each case the output looks finished, and it carries FactSet's name into work a client will act on.When AI gets something wrong in these workflows, the cost is contractual and reputational, and in some cases regulatory. The error does not stay inside our product. It travels into a client's own work product and into the decisions they make from it.AI has moved from a feature inside a few products to a capability across the platform, and our quality practice needs to scale with it. Teams have built their own test sets and their own working definitions of good, which is how most organizations start. At platform scale, conversational assistance, banker workflows, portfolio commentary, research management, and the APIs clients build on top of each need their own measures of quality, evaluated on shared tooling so the results can be compared and audited in one place.This role exists to build that function. You will define the quality and evals playbook for AI at FactSet, own the platform that measures it, and make it straightforward for every team to prove the quality of what they build. You are the first product hire into this charter, and you will build the team behind it. **What you will own****Own the evaluation platform.** You are the product leader for FactSet's shared AI evaluation platform, paired with an engineering counterpart who owns the technical build and operation. Together you stand it up and run it as enterprise infrastructure: tracing, offline and online eval harnesses, LLM-as-judge pipelines with calibrated human review, golden dataset management, and regression suites wired into CI. You own the roadmap, the adoption strategy, the integration surface into product teams' existing workflows, and the commercial and roadmap relationship with the platform vendor. Your engineering counterpart owns the technical relationship. You will know it worked when your colleagues across the enterprise run their own evals on it without your team in the loop.**Define the golden pathways.** With your engineering counterpart, develop and maintain the golden pathways for evaluation across the enterprise: the documented, opinionated way a team instruments a system, builds a dataset, runs an eval, and wires it into their release process. A product team should be able to follow the paved path without designing an evaluation approach from scratch, and you own keeping those pathways current as practice and tooling change.**Define how quality is measured.** Define the quality taxonomy for AI systems across the portfolio: what a failure is, how failures are classified, and which failure classes are non-negotiable, plus the further classes you define with product leadership and with the PMs who own each capability. This becomes the shared language the organization uses to argue about quality, and you are its steward.**Own human review operations.** Evaluation at scale depends on people producing labels on a schedule: judge calibration sets, ground-truth datasets, and adjudication of disagreements. You own that operation, including how reviewers are sourced, trained, and measured for consistency, what it costs, and how it scales as coverage grows. Your engineering counterpart builds the tooling those reviewers work in.**Set and hold the floor.** You define the minimum every AI capability must satisfy before it reaches a client. The floor is a coverage requirement: a capability ships with evals defined and running, tracing in place, a documented failure taxonomy, and a measured baseline it can be held against later. Product teams set their own quality targets above that line and are expected to aim well above it.**Drive adoption.** Enterprise platform mandates fail when the platform is slower than the spreadsheet it replaces. You will land this team by team: understand what each product group does for quality today, help them move onto shared tooling and pathways, and make the shared path the easier one. This is a sustained internal go-to-market effort and a core part of the job.**Close the loop from production.** Design the path from a client-reported failure to a permanent regression test. The program succeeds when the same failure cannot ship twice. That matters far more than the number of evals in existence.**Unite the practice.** Bring evaluation leaders at FactSet together into a small group of PMs and specialists operating as a federated center of excellence. You own the standard and the platform, and product teams own their own quality outcomes against it.**What we are looking for*** 10+ years in product management, with 4+ years on AI/ML or data-intensive products, and experience managing product managers* Demonstrated experience defining evaluation methodology for LLM or ML systems: metrics, failure taxonomies, golden datasets, and quality thresholds that engineering teams actually adopted* Experience taking an internal platform from zero to broad adoption across product teams who were not required to use it, including the migration, enablement, and change-management work that entails* Experience operating a paired product and engineering leadership model, where neither person has authority over the other's function* Hands-on fluency with modern evaluation approaches, including offline and online eval, LLM-as-judge, human-in-the-loop review, RAG and retrieval metrics, and agent trajectory evaluation, enough to design a pipeline and critique someone else's without needing to implement it* Familiarity with LLM observability and evaluation tooling such as LangSmith, LangFuse, Braintrust, or Weights & Biases. We care that you have opinions about this category rather than experience with any one product in it* A track record of standing something up rather than inheriting it. You have built a function, a standard, or a platform from nothing inside a large organization* Comfort holding a bar under commercial pressure, and the judgment to know when the bar is wrong* Ability to influence senior stakeholders across time zones without formal authority over their teams* Strong written communication. This role produces standards documents that people have to be able to follow six months later**Helpful, not required:** financial services or capital markets domain knowledge; experience with regulated or audited AI deployments; familiarity with model risk management practice.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director Product Management AI Quality And Evaluation
Director Product Management AI Quality And Evaluation

FactSet • Hyderabad

On-site
INR 3,500,000 - 6,000,000
Director, Product Management AI Quality & Evaluation
Director, Product Management AI Quality & Evaluation

FactSet • Hyderabad

On-site
INR 2,500,000 - 4,800,000
Lead Product Analyst - AI Products
Lead Product Analyst - AI Products

Praxy • India

Remote
INR 11,516,000 - 17,274,000
Director Engineering AI Evaluation And Quality
Director Engineering AI Evaluation And Quality

FactSet • Hyderabad

On-site
INR 1,200,000 - 2,400,000
AI Technical Lead
AI Technical Lead

Benchmarkit • Pune District

On-site
INR 4,000,000 - 7,000,000
Data Engineering Lead - Data Quality Systems
Data Engineering Lead - Data Quality Systems

Firmable • India

On-site
INR 3,500,000 - 6,000,000
Evals Engineer
Evals Engineer

PingAura AI Technologies Private Limited. • Mumbai

On-site
INR 1,000,000 - 1,500,000
Founding-tier ownership
Direct collaboration with AI experts
Competitive salary
Engineering Manager
Engineering Manager

Testsigma • Bengaluru

On-site
INR 7,000,000 - 12,000,000
Meaningful equity
Direct partnership with founding team
Manager, AI Engineering - Analytics
Manager, AI Engineering - Analytics

WeHireYou • India

Hybrid
INR 4,000,000 - 7,000,000
AI Product Engineer, Evidence
AI Product Engineer, Evidence

WeHireYou • India

Hybrid
INR 3,000,000 - 6,000,000