Staff Machine Learning Engineer

People In AI

New York (NY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid NYC

Job summary

People In AI seeks a Staff ML Evaluation Engineer to lead the evaluation layer for LLMs and AI systems used in high-stakes decisions. You will own offline benchmarks, golden datasets, and quality metrics tied to real outcomes, working with researchers and domain experts to translate judgment into measurable criteria.

You’ll connect offline results to live behavior, detect drift and regressions, and evolve frameworks as models and use cases change.

Qualifications

  • Own and design evaluation frameworks for AI systems used in high-stakes decisions.
  • Define quality metrics beyond accuracy for LLM-based decisions.
  • Translate intuition from researchers/domain experts into measurable criteria.

Responsibilities

  • Design and own the evaluation layer for LLMs and AI systems used in real decisions.
  • Build and maintain golden evaluation datasets grounded in expert judgment.
  • Design offline benchmarks to compare models, prompts, and retrieval strategies.
  • Define quality metrics that go beyond surface-level accuracy.
  • Partner with researchers to translate intuition into measurable criteria.
  • Connect offline evaluation results with online behavior after deployment.
  • Detect regressions, drift, and subtle failures over time.
  • Iterate on evaluation frameworks as models/use cases evolve.

Skills

Offline evaluation frameworks
Online/offline mismatch
Model monitoring
Drift detection
Regression analysis
Python
Data analysis
LLM evaluation
RAG evaluation
Prompt/system regression testing
Human-in-the-loop

Tools

Python

Job description

Staff Machine Learning Evaluation Engineer (LLMs & Decision Quality)

Hybrid NYC

We’re working with a world-class, research-driven organization operating in a high-stakes decision-making environment to hire a Staff Machine Learning Evaluation Engineer.

This role sits at the intersection of AI, data, and judgment. The focus is not on building flashy demos or optimizing infrastructure, but on answering a harder question:

When should an AI system be trusted?

The role

You’ll be responsible for designing and owning the evaluation layer for large language models and AI systems used to support real, consequential decisions.

This includes:

  • Building and maintaining golden evaluation datasets grounded in expert judgment
  • Designing offline benchmarks to compare models, prompts, and retrieval strategies
  • Defining quality metrics that go beyond surface-level accuracy
  • Partnering closely with researchers and domain experts to translate intuition into measurable criteria
  • Connecting offline evaluation results with online behavior once systems are live
  • Detecting regressions, drift, and subtle failures over time
  • Iterating on evaluation frameworks as models and use cases evolve

You’ll operate as a lead individual contributor with real influence over how AI quality is defined, measured, and enforced.

What this role is not

  • Not an ML infrastructure or serving role
  • Not a prompt-engineering or chatbot UX role
  • Not pure research with no production ownership
  • Not dashboard-only analytics

This is a hands-on engineering role focused on evaluation, trust, and decision quality.

What we’re looking for

We’re looking for someone who has actually owned model evaluation, not just consumed metrics.

Strong signals include:

  • Experience designing offline evaluation or experimentation frameworks
  • Deep understanding of online vs offline mismatch and how to manage it
  • Ownership of model monitoring, drift detection, or regression analysis
  • Comfort working with ambiguous, qualitative outputs (e.g. LLMs)
  • Experience partnering with domain experts or stakeholders to define “what good looks like”
  • Strong Python and data skills; comfort building lightweight pipelines and analysis tooling

Backgrounds that tend to work well:

  • ML tech leads
  • Research engineers with production exposure
  • Engineers from decisioning, risk, pricing, trust & safety, or ranking systems

LLM experience

Hands-on experience with LLMs is highly relevant, especially around:

  • LLM evaluation and benchmarking
  • RAG evaluation
  • Hallucination or grounding checks
  • Prompt or system regression testing
  • Human-in-the-loop review workflows

Why this role is interesting

  • You’ll shape how AI systems are measured and trusted, not just how they’re built
  • You’ll work on problems where mistakes matter
  • You’ll influence real decisions, not vanity metrics
  • You’ll have autonomy to define standards, not just follow them

Compensation & location

  • Total compensation is highly competitive and aligned with top-tier technology firms
  • Location flexible within the U.S.

If you’re excited by the idea of building the yardsticks that decide whether AI systems are actually helping or quietly harming decision-making, we’d love to hear from you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of AI Evaluation
Head of AI Evaluation

People In AI • New York (NY)

On-site
USD 180,000 - 280,000
Staff ML Evaluation Engineer - Trust in AI Decisions
Staff ML Evaluation Engineer - Trust in AI Decisions

People In AI • New York (NY)

Hybrid
USD 180,000 - 240,000
Hybrid NYC
Machine Learning Scientist
Machine Learning Scientist

Arena Intelligence, Inc. • San Francisco (CA)

On-site
USD 140,000 - 220,000
Competitive compensation
Comprehensive health and wellness benefits
Cutting-edge AI projects
+1
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
Applied Research Scientist, LLM Evaluation & Post-Training
Applied Research Scientist, LLM Evaluation & Post-Training

Innodata Inc. • United States

On-site
USD 175,000 - 225,000
Machine Learning Scientist - Open Source Lead
Machine Learning Scientist - Open Source Lead

Arena Intelligence, Inc. • San Francisco (CA)

On-site
USD 180,000 - 250,000
Competitive compensation
Comprehensive health benefits
Opportunity to work on cutting-edge AI
+1
AI Engineer - GA
AI Engineer - GA

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Applied AI Engineer
Applied AI Engineer

SherlockTalent • Miami (FL)

Hybrid
USD 120,000 - 140,000
Solid Benefits
Referral bonus of $2,500
AI Engineer - FL
AI Engineer - FL

LawPro.ai • Town of Florida (NY)

On-site
USD 140,000 - 210,000
Staff AI Software Engineer
Staff AI Software Engineer

Harnham • San Francisco (CA)

On-site
USD 150,000 - 200,000