ZP-ML-004 Applied ML -- eval & alignment Remote · US RESEARCH →

Zute Predictive

Northern (KY)

Hybrid

USD 110,000 - 160,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Zute Predictive is seeking an engineer to design evaluation harnesses that verify model performance after deployments and upgrades. You will build pipelines that run in CI and help prevent regressions before they reach production.

You will define measurable criteria for what "working" means across customer integrations, design annotation workflows to keep ground truth aligned with evolving data, and research robustness under distribution shifts and adversarial inputs.

Qualifications

  • Experience building eval infrastructure that caught real production problems.
  • Familiar with frontier models at the API level.
  • Able to maintain Python in production-ready form.
  • Understanding of statistical significance for eval results.

Responsibilities

  • Build eval pipelines that run automatically on every model upgrade and flag regressions before production.
  • Define what “working” means for each customer integration in measurable terms.
  • Design annotation workflows to keep ground truth current as data drifts.
  • Research prompt robustness: adversarial inputs, distribution shift, context window edge cases.
  • Publish internal findings as reference documents for the team.

Skills

Python
Eval pipelines
Statistical significance
Adversarial inputs

Tools

CI/CD tooling
PyTest

Job description

Design the evaluation harnesses that tell us whether a deployed model is actually working. Build them to run in CI and to survive model upgrades.

WHAT YOU'LL DO
  • Build eval pipelines that run automatically on every model upgrade and flag regressions before they reach production.
  • Define what \"working\" means for each customer integration -- in measurable, reproducible terms.
  • Design the annotation workflows that keep ground truth current as customer data drifts.
  • Research prompt robustness: adversarial inputs, distribution shift, context window edge cases.
  • Publish internal findings as reference documents the whole team can act on.
WHAT WE'RE LOOKING FOR
  • Track record building eval infrastructure that actually caught a real production problem.
  • Familiar with the major frontier models at the API level -- not just the papers.
  • Can write Python that a systems engineer would be comfortable maintaining.
  • Understand statistical significance well enough to know when an eval result is meaningful.
  • Have written a prompt that was still working six months later.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote ML Evaluation Engineer
Remote ML Evaluation Engineer

Zute Predictive • Northern (KY)

Hybrid
USD 110,000 - 160,000
ML Engineer, Post-Training
ML Engineer, Post-Training

Zoro • Northern (KY)

On-site
USD 120,000 - 180,000
Remote ML Engineer — Pipelines, Evaluation & Impact
Remote ML Engineer — Pipelines, Evaluation & Impact

Greenhouse Software, Inc. • United States

Remote
USD 120,000 - 180,000
40 paid days off
Remote work from anywhere in the world
Competitive salary and bonuses
ML Engineer
ML Engineer

Greenhouse Software, Inc. • United States

On-site
USD 120,000 - 180,000
40 paid days off
Remote work from anywhere in the world
Competitive salary and bonuses
ML Research Engineer - PhD - AI Trainer
ML Research Engineer - PhD - AI Trainer

Mercor • Seattle (WA)

On-site
USD 100,000 - 150,000
Model Evaluation Engineer
Model Evaluation Engineer

Zof AI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 190,000
MacBook Pro
Premium AI development tools
OpenAI Codex Max
Forward Deployed ML Engineer
Forward Deployed ML Engineer

Alexander Chapman • United States

On-site
USD 140,000 - 190,000
Senior Machine Learning Engineer (Generative AI)
Senior Machine Learning Engineer (Generative AI)

Kilwa • Chicago (IL), Northern (KY)

Hybrid
USD 120,000 - 180,000
Founding Forward Deployed Machine Learning Engineer [33151]
Founding Forward Deployed Machine Learning Engineer [33151]

Stealth Startup • Sunnyvale (CA)

On-site
USD 160,000 - 220,000
0.5-2.0% Equity
Insurance
Backend Software Engineer (Evals)
Backend Software Engineer (Evals)

OpenAI • Seattle (WA)

On-site
USD 230,000 - 385,000