Remote AI Eval Engineer: Ground Truth & Datasets

Entrada Ventures

United States

Hybrid

USD 120,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical, dental & vision insurance
Equity and meaningful early-stage"
Unlimited PTO
Hybrid/Remote stipend
In-office perks and snacks

Job summary

Prophetic Software is seeking an engineer to build and defend evaluation datasets for a scalable ML stack. You will own ground truth definitions, data sourcing, labeling guidelines, and the evaluation loop, ensuring trusted data drives product decisions.

You’ll work with product and engineering to decompose systems into testable modules, define correctness, and measure performance. This role emphasizes hands-on data work and cross-team collaboration.

Qualifications

  • Experience building or evaluating ML/LLM systems in production.
  • You have built evaluation datasets and can discuss one in detail.
  • Working fluency in ML validation fundamentals: train/validation/test, precision/recall, calibration, inter-rater agreement.
  • Understand feature engineering to reason about classifier failures.
  • Strong Python and SQL; able to pull and reshape data yourself.
  • Hands-on with prompting, structured outputs, and tool-use harnesses for LLMs.
  • Judgment about when LLM-as-judge is reliable and how to prove it.
  • You can read a system design and translate it into a label schema.

Responsibilities

  • Decide what needs to be measured and how.
  • Define correctness with rubrics, label schemas, and tolerances.
  • Prioritize validation sets to unblock iterations.
  • Source data via production sampling, labeling programs, and synthetic ground truth.
  • Manage labeling vendors or internal experts and ensure label quality.
  • Calibrate automated graders against human gold sets.
  • Report metrics like precision, recall, F1, and calibration.
  • Collaborate with engineers to integrate datasets into pipelines.

Skills

ML/LLM systems in production
Evaluation datasets
Python
SQL
ML validation fundamentals
Labeling/vendor management
Data sampling
System design translation

Job description

Prophetic Software is seeking an engineer to build and defend evaluation datasets for a scalable ML stack. You will own ground truth definitions, data sourcing, labeling guidelines, and the evaluation loop, ensuring trusted data drives product decisions.

You’ll work with product and engineering to decompose systems into testable modules, define correctness, and measure performance. This role emphasizes hands-on data work and cross-team collaboration.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Evaluation Data Engineer — Ground Truth
Remote AI Evaluation Data Engineer — Ground Truth

Propheticsoftware • Northern (KY)

Hybrid
USD 120,000 - 190,000
Remote AI Software Engineer - Code & Data Evaluation
Remote AI Software Engineer - Code & Data Evaluation

Turing • United States

Remote
USD 83,000 - 138,000
Remote Model Evaluation & Validation Engineer
Remote Model Evaluation & Validation Engineer

DEFCON AI, Inc. • United States

Remote
USD 150,000 - 190,000
Fully remote role
Equity package
Health insurance for you and family
+3
Remote AI Software Engineer – ML Data & Code Evaluation
Remote AI Software Engineer – ML Data & Code Evaluation

engineeringjobs.net, Inc. • Pittsburgh

Remote
USD 90,000 - 130,000
Staff Engineer - AI Evaluation & Metrics Platform
Staff Engineer - AI Evaluation & Metrics Platform

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 200,000
Remote AI Evaluation Architect - Tech Docs & Code
Remote AI Evaluation Architect - Tech Docs & Code

Weekday 1 • United States

Remote
USD 21,000 - 28,000
Fully remote
Flexible hours
Weekly payments
Remote AI Software Engineer: Code Quality & Evaluation
Remote AI Software Engineer: Code Quality & Evaluation

Turing • United States

Remote
USD 55,000 - 110,000
Remote AI Benchmark & Datasets Engineer
Remote AI Benchmark & Datasets Engineer

Pathway Genomics Corporation • Palo Alto (CA)

Remote
USD 150,000 - 210,000
Intellectually stimulating work environment
Work with a pioneering AI startup
Flexible remote work options
Forward Deployed ML Engineer
Forward Deployed ML Engineer

Alexander Chapman • United States

On-site
USD 140,000 - 190,000
Research Engineer – AI Evaluation & Metrics
Research Engineer – AI Evaluation & Metrics

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000