Quality Lead, Agentic AI Workflow Evaluation

Innodata

San Jose (CA)

On-site

USD 103,000 - 117,000

Full time

13 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Innodata is seeking a Quality Lead to stand up an onsite team in San Jose for evaluating complex agentic AI workflows. You’ll own the audit system, calibrate reviewer output, and drive improvements with the customer.

The role is a senior individual contributor, with responsibility for defining scoring standards, maintaining rubrics, and training reviewers. You will report insights to the Engagement Manager and customer while ensuring privacy and security practices onsite.

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 4+ years in quality assurance or senior review roles in annotation, evaluation, trust and safety, or a judgment-intensive domain.
  • Experience owning a quality function: designed the audit, not just executed.
  • Significant experience with AI/ML evaluation: annotation, red-teaming, RLHF, model/agent evaluation.
  • Hands-on familiarity with agentic systems: tool use, multi-step task execution, sandboxed environments.
  • Ability to run calibration with peers and resolve disagreements.
  • Strong written communication to document scoring standards clearly.
  • Proficiency with spreadsheets and dashboarding tools; some Python/SQL for data analysis.
  • Experience training/reviewers into rubric-based work.

Responsibilities

  • Own the quality system for the engagement: audit design, sampling strategy, scoring standards, and reporting.
  • Build the system with the customer's quality leads and drive concrete improvements.
  • Re-score a sample of reviewer output as a second pass and identify error patterns.
  • Run calibration sessions and document the reasoning for future consistency.
  • Maintain rubric health and drive revisions with the customer.
  • Train and onboard new reviewers, including ramp criteria and readiness checks.
  • Provide evidence behind performance conversations to the Engagement Manager.
  • Report quality trends to the Engagement Manager and the customer.
  • Deputize for the Engagement Manager on delivery operations during absences.
  • Maintain information security, privacy, and facility access practices for onsite environment.

Skills

QA governance
AI evaluation
Calibration sessions
Documentation
Spreadsheet basics
Python/SQL basics

Education

Bachelor's degree or equivalent practical experience

Job description

Quality Lead, Agentic AI Workflow Evaluation

In Office - San Jose, California

Innodata(Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked.Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale.We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role:

We are standing up a dedicated onsite team to evaluate complex, real-world agentic AI workflows for a frontier AI customer. Reviewers work through ambiguous, multi-step scenarios inside isolated test environments, assessing whether AI agents complete tasks safely, respect user intent and consent, and hold up under close scrutiny. The Quality Lead is the person accountable for whether that output is any good.

This is a senior individual contributor role. You will not manage the reviewers — that sits with the Engagement Manager — but you set the standard they are held to. You own the audit sample, run calibration, keep the rubric usable as real cases stress it, and train reviewers into the work. You are also the deputy: when the Engagement Manager is out, the engagement runs on you.

The quality approach here is not fully defined. We expect you to build it in partnership with the customer's quality leads, or at minimum to take what they have, run it honestly, and come back with specific recommendations for where it falls short.

What You’ll Own:
  • Own the quality system for the engagement: audit design, sampling strategy, scoring standards, and how quality gets measured and reported
  • Build that system with the customer's quality leads where none exists, and where one does, operate it and recommend concrete improvements based on what the data shows
  • Re-score a sample of reviewer output as a second pass; identify error patterns rather than isolated mistakes
  • Run calibration sessions: surface disagreement, work it to resolution, and document the reasoning so the outcome holds for future cases
  • Maintain rubric health — flag criteria that are ambiguous, overlapping, or silent on cases the team keeps hitting, and drive revisions through the customer
  • Train and onboard new reviewers, including nesting plans, ramp criteria, and the judgment call on when someone is production-ready
  • Give the Engagement Manager the evidence behind performance conversations: who is drifting, on what, and whether coaching is working
  • Report quality trends to the Engagement Manager and, alongside them, to the customer
  • Deputize for the Engagement Manager on delivery operations during absences
  • Maintain information security, privacy, and facility access practices required by the customer's onsite environment
You’ll Thrive in This Role If You Have:
  • Bachelor's degree or equivalent practical experience
  • 4+ years in quality assurance, quality management, or senior review work within annotation, evaluation, trust and safety, or a similarly judgment-intensive domain
  • Direct experience owning a quality function: you designed the audit, not just executed someone else’s
  • Significant experience with AI/ML evaluation work: annotation, red-teaming, RLHF, model or agent evaluation, or trust and safety review
  • Hands-on familiarity with agentic systems: tool use, multi-step task execution, sandboxed environments, and common failure modes
  • Demonstrated ability to run calibration with peers — including holding a position under disagreement and changing it when the argument is better
  • Strong written communication; able to document a scoring standard clearly enough that a reviewer can apply it and an auditor can check it
  • Comfortable in spreadsheets and in a dashboarding tool, with enough Python or SQL to pull and slice your own data (you will not be asked to build interfaces)
  • Experience training or onboarding reviewers into rubric-based work

The expected hourly salary range for this position is $75-85 p/hour, based on experience, skills, and qualifications.

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams.

As set forth in Innodata Inc.’s Equal Employment Opportunity policy,we do not discriminate on the basis of any protected group status under any applicable law.

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service‑connected disability.

A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Quality Lead, Agentic AI Workflow Evaluation
Quality Lead, Agentic AI Workflow Evaluation

Innodata Inc. • San Jose (CA)

On-site
USD 103,000 - 117,000
Engagement Manager, Agentic AI Workflow Evaluations
Engagement Manager, Agentic AI Workflow Evaluations

Innodata • San Jose (CA)

On-site
USD 103,000 - 117,000
Technical Training & Quality Manager
Technical Training & Quality Manager

Innodata • Austin (TX)

On-site
USD 145,000 - 175,000
Data AI Annotator
Data AI Annotator

Innodata Inc. • Northern (KY)

Hybrid
USD 18,000 - 23,000
Technical Delivery Manager
Technical Delivery Manager

Innodata • New York (NY)

On-site
USD 145,000 - 165,000
Senior Quality Lead — Agentic AI Workflow Evaluation
Senior Quality Lead — Agentic AI Workflow Evaluation

Innodata Inc. • San Jose (CA)

On-site
USD 103,000 - 117,000
Senior QA Lead, Agentic AI Workflows
Senior QA Lead, Agentic AI Workflows

Innodata • San Jose (CA)

On-site
USD 103,000 - 117,000
Sales Development Representative
Sales Development Representative

Innodata • Ridgefield Park (NJ)

On-site
USD 80,000 - 100,000
Senior AI Platform Engineer
Senior AI Platform Engineer

EarnIn • Mountain View (CA)

On-site
USD 228,000 - 279,000
Head of AI Quality, Assurance and Evaluation
Head of AI Quality, Assurance and Evaluation

Agilent Technologies • Santa Clara (CA)

On-site
USD 180,000 - 260,000