AI Evaluation Engineer

Capital Rx

Denver (CO)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Judi Health is seeking an AI Evaluation Engineer to build testing frameworks, metrics, and tooling to assess the safety, reliability, and accuracy of AI models in production. You will translate ambiguous product goals into measurable quality targets and drive standardized evaluation across multiple agent types.

You will own data engineering tasks, platform observability, and evaluation tooling, collaborating with data science and engineering teams to ensure high-quality, reproducible results and

Qualifications

  • 4+ years of experience in data engineering, ML engineering, or software engineering.
  • Bachelor's or Master's degree in Computer Science, Machine Learning, or a related quantitative field.
  • Strong proficiency in Python.
  • Strong SQL skills.
  • Experience building and maintaining production data pipelines.
  • Experience working with at least one cloud platform (AWS preferred).

Responsibilities

  • Data Engineering: Build and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback).
  • Platform & Observability: Develop dashboards and monitoring tools for AI quality metrics and integrate evaluations into CI/CD pipelines.
  • AI / ML Evaluation Tooling: Apply and extend evaluation patterns and design metrics for stochastic systems using LangSmith and related tools.
  • Collaboration: Partner with data science, engineering, and product teams to translate goals into engineering requirements and ensure high usability of tooling.

Skills

Python
SQL
Cloud platforms (AWS)
Data pipelines

Education

Bachelor's or Master's in CS/ML/related field

Tools

LangSmith
Playwright

Job description

Position Summary

As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development and real‑world usage by translating ambiguous product goals into measurable quality targets.

What You’ll Build
  • Evaluation & Quality Pipelines
    • Build data evaluation pipelines that collect production conversations and agent interactions.
    • Reconstruct full sessions from traces, logs, recordings, and transcripts.
    • Apply labeling and scoring using human feedback signals (surveys, sentiment, outcomes) and automated evaluators (e.g., LLMasjudge).
  • Continuous Quality & Safety Benchmarking
    • Own weekly and on‑demand automated evaluation runs against staging and production.
    • Define benchmarks that track accuracy, reliability, and safety‑related signals.
    • Produce trend dashboards that clearly answer: “Did this deploy change quality or risk?”
  • Unified Evaluation Framework
    • Design and extend a standardized evaluation framework that supports multiple agent types and workflows.
    • Translate high‑level product expectations into concrete success criteria and metrics.
    • Ensure new agents and features can be evaluated consistently with minimal friction.
  • Self‑Service Evaluation Tooling
    • Build APIs and internal tools so data scientists and engineers can go from “interesting scenario” to “included in the eval suite” quickly.
    • Enable scenario curation, dataset management, and eval execution without deep infrastructure knowledge.
  • Experiment Tracking & Visibility
    • Provide shared visibility into prompt, model, and agent experiments.
    • Enable reproducibility and comparison across runs so teams can build on each other’s work instead of operating in silos.
Position Responsibilities
  • Data Engineering
    • Build and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback).
    • Implement complex data stitching and session reconstruction logic.
    • Manage dataset versioning, provenance, and lifecycle.
  • Platform & Observability
    • Develop dashboards and monitoring tools for AI quality metrics.
    • Integrate evaluations into CI/CD pipelines for scheduled and gated runs.
    • Implement alerting on quality and safety signals, not just infrastructure health.
  • AI / ML Evaluation Tooling
    • Apply and extend LLMasjudge evaluation patterns.
    • Design metrics and scoring approaches suitable for stochastic, nondeterministic systems.
    • Use tools like LangSmith to track runs, traces, experiments, and evaluation results.
  • Collaboration
    • Partner closely with data science, engineering, and product teams.
    • Translate between research goals, product intent, and engineering constraints.
    • Help define what “good” looks like for AI behavior in production.
    • Advocate for strong developer experience and usability in the tools you build.
    • Responsible for adherence to the Capital Rx Code of Conduct including the reporting of non‑compliance.
Required Qualifications
  • 4+ years of experience in data engineering, ML engineering, or software engineering.
  • Bachelor's or Master's degree in Computer Science, Machine Learning, or a related quantitative field.
  • Strong proficiency in Python.
  • Experience building and maintaining production data pipelines.
  • Strong SQL skills.
  • Experience working with at least one cloud platform (AWS preferred).
Nice to Haves
  • Prior work on LLM or agent evaluation infrastructure.
  • Familiarity with designing metrics for safety, reliability, or quality in AI systems.
  • Experience with voice or call center data (audio, transcripts, sentiment).
  • Experience with browser automation tools (e.g., Playwright) for end‑to‑end evals.
  • Deep SQL expertise.

All employees are responsible for adherence to the Judi Health Code of Conduct including the reporting of non‑compliance. This position description is designed to be flexible, allowing management the opportunity to assign or re‑assign duties and responsibilities as needed to best meet organizational goals.

We provide equal employment opportunities to all employees and applicants for employment and prohibit discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, medical condition, genetic information, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer
AI Evaluation Engineer

Capital Rx • Charlotte (NC)

On-site
USD 120,000 - 180,000
AI Evaluation Engineer
AI Evaluation Engineer

Capital Rx • New York (NY)

On-site
USD 120,000 - 160,000
AI Evaluation Engineer
AI Evaluation Engineer

Transformcap • United States

On-site
USD 135,000 - 200,000
AI Evaluation Engineer
AI Evaluation Engineer

Capitalrx • Charlotte (NC)

On-site
USD 134,800 - 168,500
Senior QA Engineer
Senior QA Engineer

Third Way Health • Cambridge (MA)

On-site
USD 130,000 - 180,000
AI Quality & Safety Evaluation Engineer
AI Quality & Safety Evaluation Engineer

Capital Rx • New York (NY)

On-site
USD 120,000 - 160,000
Senior Backend Engineer, AI Evaluations
Senior Backend Engineer, AI Evaluations

Ellipsis Health • San Francisco (CA)

On-site
USD 160,000 - 210,000
401(k) matching
Health insurance
Vision insurance
+2
Senior QA Engineer
Senior QA Engineer

Third Way Health, Inc. • Cambridge (MA)

On-site
USD 120,000 - 150,000
Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Spain (TX)

On-site
EUR 70,000 - 100,000
Evals Lead
Evals Lead

Fluency Digital, Inc. • New York (NY)

On-site
USD 120,000 - 150,000