Evals Lead

Fluency Digital, Inc.

New York (NY)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fluency Digital, Inc. in New York is seeking an Evals Lead to own the evaluation frameworks for AI agents managing clinical administrative tasks. This role involves designing evaluation pipelines, ensuring that models are reliable and safe before deployment in patient workflows.

The ideal candidate will have over 5 years of experience, particularly with ML evaluations, strong SQL and Python skills, and should be based in the United States.

Qualifications

  • 5+ years in QA, test engineering, or ML evaluation, with 2 years in LLM systems.
  • Strong SQL and Python skills; comfortable with MongoDB and PostgreSQL.
  • Experience designing automated regression suites and qualitative eval rubrics.

Responsibilities

  • Own frameworks for evaluating models in production workflows.
  • Design pipelines that blend correctness checks with evaluative rubrics.
  • Define healthcare workflow metrics to ensure clinical safety.

Skills

SQL
Python
ML evaluation
Communication

Tools

GCP
MongoDB
PostgreSQL

Job description

AI company building the invisible layer that handles clinical administrative work so providers don't have to

We're building agents that handle the phone calls, faxes, prior auths, and scheduling loops that quietly eat half a clinic's day. As Evals Lead, you own the frameworks that tell us whether those agents are actually good enough to deploy into live patient workflows—where a hallucinated prior auth status isn't a benchmark metric, it's a real person waiting for care. You'll design evaluation pipelines that blend deterministic correctness checks with LLM-as-judge rubrics, surface failure patterns before they reach production, and build the data flywheel that makes our models measurably safer and more reliable every week.

What we're looking for:
  • 5+ years in QA, test engineering, or ML evaluation, with at least 2 years focused on evaluating LLM-based systems or conversational AI in production
  • Strong SQL and Python skills; comfortable working across MongoDB, PostgreSQL, and GCP to instrument evaluation infrastructure
  • Experience designing both automated regression suites and qualitative eval rubrics that capture task completion, tone, and clinical safety
  • Ability to define metrics that matter for healthcare workflows—high recall on prior auth submission errors matters more than abstract accuracy
  • Clear communication instincts; you can explain why a model failed to engineers and product stakeholders without hand-waving
  • Must be based in the United States.
Tech stack:

GCP, Mongo, MongoDB, PostgreSQL

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior QA Engineer
Senior QA Engineer

Third Way Health, Inc. • Cambridge (MA)

On-site
USD 120,000 - 150,000
Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Spain (TX)

On-site
EUR 70,000 - 100,000
AI Evaluation Engineer
AI Evaluation Engineer

Capitalrx • Charlotte (NC)

On-site
USD 134,800 - 168,500
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
Software Engineer, AI Evaluation
Software Engineer, AI Evaluation

Nuna Inc. • San Francisco (CA)

On-site
USD 180,000 - 270,000
AI Engineer - VA
AI Engineer - VA

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Town of Poland (NY)

On-site
USD 140,000 - 200,000
Clinical Research Scientist, Mental Health AI
Clinical Research Scientist, Mental Health AI

Vals AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Significant ownership over research
Collaboration with clinical and tech
Equity and title flexibility
+4
AI Engineer - TX
AI Engineer - TX

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000