Get more replies from employers
Send a job-specific resume in minutes.
Fluency Digital, Inc. in New York is seeking an Evals Lead to own the evaluation frameworks for AI agents managing clinical administrative tasks. This role involves designing evaluation pipelines, ensuring that models are reliable and safe before deployment in patient workflows.
The ideal candidate will have over 5 years of experience, particularly with ML evaluations, strong SQL and Python skills, and should be based in the United States.
AI company building the invisible layer that handles clinical administrative work so providers don't have to
We're building agents that handle the phone calls, faxes, prior auths, and scheduling loops that quietly eat half a clinic's day. As Evals Lead, you own the frameworks that tell us whether those agents are actually good enough to deploy into live patient workflows—where a hallucinated prior auth status isn't a benchmark metric, it's a real person waiting for care. You'll design evaluation pipelines that blend deterministic correctness checks with LLM-as-judge rubrics, surface failure patterns before they reach production, and build the data flywheel that makes our models measurably safer and more reliable every week.
GCP, Mongo, MongoDB, PostgreSQL