Data Scientist — Agent Evaluations & Quality

Clera

United States

Remote

USD 120,000 - 190,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Clera is building an AI executive assistant operating across email, calendars, meetings, and business software. We seek a Data Scientist — Agent Evaluations & Quality to own the measurement system that determines real-world improvements in ambiguous environments.

You will partner with AI Agent Capabilities engineers to generate evidence shaping product decisions and release quality. This high-ownership role sits at the intersection of applied data science, LLM evaluation, and product quality,

Qualifications

  • 4+ years in Applied Data Science or Machine Learning roles, with a track record of building and delivering evaluation systems, automated data pipelines, or production ML infrastructure.
  • Experience designing automated evaluation frameworks, success criteria, and regression suites for complex AI/ML or agentic systems.
  • Production-grade proficiency in Python and SQL, with experience building and maintaining automated analytical pipelines on large datasets.
  • Applied statistical and experimental skills: significance testing, variance analysis, and sampling to evaluate non-deterministic AI/ML systems.
  • Experience developing labeled datasets, annotation guidelines, and quality-control processes for ground-truth data in dynamic product environments.
  • Solid understanding of LLM agent behaviors: tool use, multi-step execution, retrieval, and practical failure modes.
  • Demonstrated ability to analyze model traces, tool calls, and outputs to identify root causes across model, prompt, tool, and data layers.
  • Experience using production telemetry and observability data to monitor system quality, build dashboards, and analyze real-world user outcomes.

Responsibilities

  • Architect and maintain automated evaluation pipelines that measure agent quality across product surfaces.
  • Translate agent capabilities into explicit pass, partial-pass, and failure criteria for complex multi-step tasks.
  • Build representative gold datasets and regression suites covering real workflows, edge cases, and adversarial scenarios.
  • Define meaningful metrics — task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.
  • Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and track grader agreement.
  • Compare models, prompts, and implementations using rigorous offline experiments and production evidence.
  • Analyze traces and production outcomes to identify root causes and build a practical failure taxonomy.
  • Turn production failures into regression cases and continuously close gaps in evaluation coverage.
  • Build dashboards and release-quality signals that make results actionable for engineering, product, and leadership.
  • Recommend improvements to capability engineers and verify that fixes raise quality without unacceptable regressions.

Skills

Python
SQL
Evaluation pipelines
Experiment design
Statistical analysis
LLM evaluation

Job description

About the Role

This company is building an AI executive assistant that operates across email, calendars, meetings, and business software. As a Data Scientist — Agent Evaluations & Quality, you will own the measurement system that determines whether the assistant is genuinely improving in ambiguous, real-world environments. You'll partner directly with AI Agent Capabilities engineers to generate the evidence that shapes product decisions, model choices, and release quality.

This is a high-ownership, deeply technical role at the intersection of applied data science, LLM evaluation, and product quality — ideal for someone who thrives on turning hard, open-ended quality questions into rigorous, actionable answers.

What You'll Do
  • Architect and maintain automated evaluation pipelines that measure agent quality across product surfaces.

  • Translate agent capabilities into explicit pass, partial-pass, and failure criteria for complex multi-step tasks.

  • Build representative gold datasets and regression suites covering real workflows, edge cases, and adversarial scenarios.

  • Define meaningful metrics — task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.

  • Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and track grader agreement.

  • Compare models, prompts, and implementations using rigorous offline experiments and production evidence.

  • Analyze traces and production outcomes to identify root causes and build a practical failure taxonomy.

  • Turn production failures into regression cases and continuously close gaps in evaluation coverage.

  • Build dashboards and release-quality signals that make results actionable for engineering, product, and leadership.

  • Recommend improvements to capability engineers and verify that fixes raise quality without unacceptable regressions.

What We're Looking For

Required

  • 4+ years in Applied Data Science or Machine Learning roles, with a track record of building and delivering evaluation systems, automated data pipelines, or production ML infrastructure.

  • Experience designing and implementing automated evaluation frameworks, success criteria, and regression suites for complex AI/ML or agentic systems.

  • Production-grade proficiency in Python and SQL, with experience building and maintaining automated analytical pipelines on large datasets.

  • Applied statistical and experimental skills: significance testing, variance analysis, and sampling to evaluate non-deterministic AI/ML systems.

  • Experience developing labeled datasets, annotation guidelines, and quality-control processes for ground-truth data in dynamic product environments.

  • Solid understanding of LLM agent behaviors: tool use, multi-step execution, retrieval, and practical failure modes.

  • Demonstrated ability to analyze model traces, tool calls, and outputs to identify root causes across model, prompt, tool, and data layers.

  • Experience using production telemetry and observability data to monitor system quality, build dashboards, and analyze real-world user outcomes.

Nice to Have

  • Hands-on experience with LLM-as-a-judge systems, model-based grading, or AI benchmarking platforms.

  • Experience shipping or operating production ML products, agentic systems, or customer-facing consumer software.

  • Experience reviewing and adapting public research benchmarks or academic evaluation methodologies to real-world product problems.

What makes you a great fit

  • You're product-oriented — you prioritize metrics tied to real user outcomes, not just convenient measurements.

  • You drive ambiguous quality questions from evaluation design all the way into product decisions.

  • You write maintainable, production-quality code — not just ad-hoc notebooks.

  • You collaborate naturally with engineers and are comfortable digging into traces and system internals.

Location

This role is on-site. Visa sponsorship is not available for this position.

Compensation & Benefits

Compensation details were not provided for this listing. A competitive package commensurate with experience is expected at this stage of company growth.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Agents
Member of Technical Staff, Agents

Engg • Palo Alto (CA)

On-site
USD 200,000 - 300,000
QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
Data Scientist
Data Scientist

ScienceLogic • United States

On-site
USD 120,000 - 180,000
QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners, LLC • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners • Town of Florida (NY), Northern (KY)

Hybrid
USD 120,000 - 150,000
Data Scientist, AI Agent Quality & Evaluation
Data Scientist, AI Agent Quality & Evaluation

Clera • United States

Remote
USD 120,000 - 190,000
QA Engineer - Agentic Systems
QA Engineer - Agentic Systems

Meet Life Sciences • New York (NY)

On-site
USD 110,000 - 170,000
100% Remote – QA Automation OR Data Scientist with AI Exp.
100% Remote – QA Automation OR Data Scientist with AI Exp.

SDH Systems • United States

Remote
USD 120,000 - 155,000
Data Scientist (AI Quality & Evaluation)
Data Scientist (AI Quality & Evaluation)

Bioscope.ai, Inc. • Boston (MA)

On-site
USD 100,000 - 130,000
ML Research Engineer
ML Research Engineer

Radical AI • New York (NY)

On-site
USD 140,000 - 210,000