AI Reliability Lead: Build Trustworthy, High-Quality AI

PowerToFly

Eagan (MN)

Hybrid

USD 91,000 - 169,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Hybrid work model
Flexible vacation
Mental health days

Job summary

Thomson Reuters in the United States is seeking an AI Reliability Manager to lead evaluation and quality measurement for agentic AI within the CLEAR investigative platform. This hands‑on role develops gold datasets, annotation guidelines, and scoring rubrics, translating findings into actionable engineering work and release readiness.

You will partner with product, engineering, and data science teams to diagnose failure patterns, guide improvements, and support cross‑functional annotators

Qualifications

  • Three or more years, or equivalent experience, in work requiring careful judgment about quality, such as research, analysis, editorial, audit, quality assurance, or investigative work
  • Experience applying consistent standards to open‑ended work where reasonable people can disagree about what "good" looks like
  • Hands‑on use of generative AI tools, with enough curiosity to have noticed how they fail, including confident answers that aren't supported and sources that don't say what the AI claims
  • Strong written communication, including explaining technical findings to non‑technical audiences

Responsibilities

  • Own and extend the gold datasets, annotation guidelines, and scoring rubrics used to evaluate agentic AI output for accuracy, completeness, and appropriate sourcing
  • Execute evaluation cycles, scoring AI responses against established guidelines, maintaining consistency across annotators, and flagging where output falls short
  • Analyze evaluation results to identify failure patterns, quantify their impact, and recommend where engineering and data science effort should be focused
  • Test and validate AI features prior to release, and provide a clear, evidence-based point of view on release readiness
  • Train and support additional cross functional annotator resources, including globally distributed contributors who participate in evaluation cycles but are not day‑to‑day experts, ensuring they apply guidelines consistently
  • Serve as a day‑to‑day product resource for the sales channel, answering questions about product behavior, capabilities, and known limitations, and keeping them current on the status of open issues
  • Own intake of reported issues across the CLEAR application, from AI output quality concerns to general product defects; reproduce issues, determine root cause category, and document them precisely
  • Open and manage engineering and labs tickets, and drive them to resolution alongside product, engineering, and data science partners
  • Track recurring defect themes over time and surface them to product leadership as systemic issues rather than one‑off tickets
  • Partner with product management, engineering, applied research, and go‑to‑market teams to translate quality findings into roadmap and prioritization decisions
  • Build an understanding of customer pain points and goals, and help identify new ways agentic AI can be applied to investigative workflows

Skills

Quality judgment
AI tooling
Written communication

Tools

Generative AI tools

Job description

Thomson Reuters in the United States is seeking an AI Reliability Manager to lead evaluation and quality measurement for agentic AI within the CLEAR investigative platform. This hands‑on role develops gold datasets, annotation guidelines, and scoring rubrics, translating findings into actionable engineering work and release readiness.

You will partner with product, engineering, and data science teams to diagnose failure patterns, guide improvements, and support cross‑functional annotators

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Reliability & Quality Lead
AI Reliability & Quality Lead

Thomson Reuters • Eagan (MN)

Hybrid
USD 91,000 - 169,000
Hybrid Work Model
Mental Health Days
Lead AI Engineer — Enterprise Systems & RAG Architect
Lead AI Engineer — Enterprise Systems & RAG Architect

Thomson Reuters • Eagan (MN)

Hybrid
USD 127,400 - 236,600
Hybrid Work Model
Competitive Benefits
Mental Health Days
Lead AI Engineer — Enterprise-Scale AI & Agents
Lead AI Engineer — Enterprise-Scale AI & Agents

Reuters • Minnesota

Hybrid
USD 127,400 - 236,600
Hybrid work model
Competitive benefits
Tuition reimbursement
AI Reliability Engineering Leader for Production‑Scale AI
AI Reliability Engineering Leader for Production‑Scale AI

84.51 Centre • Cincinnati (OH)

On-site
USD 180,000 - 240,000
AI Reliability Manager
AI Reliability Manager

PowerToFly • Eagan (MN)

On-site
USD 91,000 - 169,000
Hybrid work model
Flexible vacation
Mental health days
Senior AI Software Engineer — Full-Stack & RAG
Senior AI Software Engineer — Full-Stack & RAG

Thomson Reuters • Ann Arbor (MI)

Hybrid
USD 110,000 - 204,000
Life insurance
Parental leave
Paid time off
+3
Lead AI Engineer - Hybrid, Production AI Systems
Lead AI Engineer - Hybrid, Production AI Systems

Thomson Reuters group • Eagan (MN), Northern (KY)

Hybrid
USD 127,000 - 237,000
Hybrid Work Model
Annual Bonus
Mental Health Days
Senior AI Software Engineer — RAG & Agents (Hybrid)
Senior AI Software Engineer — RAG & Agents (Hybrid)

Thomson Reuters • Eagan (MN), Northern (KY)

Hybrid
USD 110,000 - 204,000
Hybrid work model
Annual bonus eligibility
Comprehensive benefits package
AI Reliability Engineering Leader — Production-Grade AI
AI Reliability Engineering Leader — Production-Grade AI

84.51˚ • Cincinnati (OH)

On-site
USD 190,000 - 270,000
Applied AI Science Leader — Reliable, High-Impact Systems
Applied AI Science Leader — Reliable, High-Impact Systems

Relativity • Kansas

Hybrid
USD 208,000 - 312,000