ML Engineer, Post-Training

Zoro

Northern (KY)

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Zoro, LLC is seeking a candidate to own the eval harnesses, the fine-tuning runs, and the routing layer that decides which model sees which task. You will define the judgment boundary between model and human and monitor production behavior to reduce failure rates while maintaining quality.

The role emphasizes hands-on evaluation, model routing, and continuous improvement of production accuracy and cost-efficiency.

Qualifications

  • Experience in evaluating language models and setting evaluation criteria.
  • Proficient in Python and data analysis for model evaluation.
  • Experience with fine-tuning methods (SFT, preference tuning) and deployment considerations.

Responsibilities

  • Own the eval suites that gate model changes before production.
  • Fine-tune and route models for drafting, extraction, and routing tasks.
  • Define and enforce the judgment boundary between model and human.
  • Observe production behavior and drive failure rates down week over week.
  • Lower cost per correct action without lowering the standard.

Skills

LLM evaluation
Python
Model fine-tuning
Production systems
Experimentation

Tools

Python tooling

Job description

The agents Zoro installs draft replies, route money, extract records, and decide what a human must approve. When one of them is wrong, a real company feels it. Post-training is where that behavior gets owned.

You will own the eval harnesses, the fine-tuning runs, and the routing layer that decides which model sees which task. The hardest part of the job is the judgment boundary: what the model may do alone, and what it must hand to a person.

What you will do
  • Own the eval suites that gate every model change before production
  • Fine-tune and route models for drafting, extraction, and routing tasks
  • Define and enforce the judgment boundary between model and human
  • Watch production behavior and drive failure rates down week over week
  • Lower cost per correct action without lowering the standard
What we are looking for
  • LLM systems shipped to production, not notebooks and demos
  • Deep eval methodology: you can say exactly why a model got better
  • Fine-tuning experience: SFT, preference tuning, or both
  • Strong Python and strong opinions about measurement
  • You read papers and ship code, in that order
Nice to have
  • Open-weight model deployments under real latency and cost limits
  • Retrieval systems over messy company data
  • Observability tooling for agent behavior in production
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Agent Eval Platform
Staff Software Engineer, Agent Eval Platform

Servicenow • Santa Clara (CA)

On-site
USD 180,000 - 320,000
Data Engineer - Oracle to PostgreSQL Re-Platform ( 102-08SENG-02 )
Data Engineer - Oracle to PostgreSQL Re-Platform ( 102-08SENG-02 )

OpsBrasil Serviços Cloud LTDA • United States

Remote
USD 120,000 - 180,000
Applied Scientist
Applied Scientist

Vecna AI • Chicago (IL)

On-site
USD 140,000 - 210,000
Evals Lead
Evals Lead

Aslan • Washington

On-site
USD 120,000 - 180,000
ML Engineer, Post Training
ML Engineer, Post Training

Varick Agents LTD. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Equity
Flexible PTO
Free lunch & dinner in office
+2
MTS, Post-Training (Enterprise)
MTS, Post-Training (Enterprise)

Bespoke Labs • United States

Hybrid
USD 300,000 - 350,000
Health/dental/vision coverage
401(k) plan
Daily onsite lunch
+2
Production ML Engineer: Evaluation, Fine-Tuning & Task Routing
Production ML Engineer: Evaluation, Fine-Tuning & Task Routing

Zoro • Northern (KY)

Hybrid
USD 120,000 - 180,000
Member of Technical Staff, Model Routing
Member of Technical Staff, Model Routing

Dipp AI Technologies, Inc. • New York (NY)

On-site
USD 180,000 - 230,000
Equity
Health benefits
401(k) match
+3
Machine Learning Engineer
Machine Learning Engineer

ESB Technologies • United States

On-site
USD 150,000 - 210,000
Member of Technical Staff: Research
Member of Technical Staff: Research

Zep AI • San Francisco (CA)

On-site
USD 180,000 - 260,000