LLM Evaluation Engineer - Benchmarks & Failure Analysis

Nous Research

United States

Remote

USD 130,000 - 170,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Nous Research is seeking an engineer to strengthen our evaluation systems across lab work, benchmark design, and judge calibration. You will ship evaluation infrastructure that researchers rely on from day one, while owning end-to-end eval pipelines and failure analysis.

The role emphasizes high ownership on a small team, with opportunities to shape prompts, environments, and automated graders across GAIA-like benchmarks and new tasks.

Qualifications

  • 3+ years in software/ML engineering or similar with concrete evaluation experience.
  • Experience with at least one LLM evaluation framework (Harbor, Nemo Evaluator, etc.).
  • Hands-on experience with LLMs: prompting, few-shot design, and fine-tuning or RAG.
  • Solid Python and clean, tested, version-controlled code.
  • Comfort with Git, CI/CD basics, Docker, and the Linux command line.
  • Understanding of eval statistics: kappa, confidence intervals, etc.
  • Experience with multiple agent benchmarks and adversarial eval is a plus.

Responsibilities

  • Run the full eval pipeline end to end and reproduce known results during onboarding.
  • Build a judge calibration protocol and document it for re-runs.
  • Extend benchmarks with new tasks and accompanying rubric and grader.
  • Run failure analysis on model outputs and recommend improvements.
  • Own a recurring eval workflow and ship tooling researchers actually use.

Skills

LLM evaluation
Prompting design
Python
Data analysis
Communication
Requirements planning

Tools

Git
CI/CD
Docker
Linux CLI
SSH
tmux

Job description

Nous Research is seeking an engineer to strengthen our evaluation systems across lab work, benchmark design, and judge calibration. You will ship evaluation infrastructure that researchers rely on from day one, while owning end-to-end eval pipelines and failure analysis.

The role emphasizes high ownership on a small team, with opportunities to shape prompts, environments, and automated graders across GAIA-like benchmarks and new tasks.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

On-site
USD 166,000 - 207,000
Health coverage
Equity
Retirement benefits
+3
LLM Benchmark Engineer — Elevate Model Evaluation
LLM Benchmark Engineer — Elevate Model Evaluation

Vals AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Relocation support
Health insurance
Lunch provided
+3
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
GenAI Evaluation Scientist - LLM Benchmarks & Failures
GenAI Evaluation Scientist - LLM Benchmarks & Failures

Scale • San Francisco (CA), Seattle (WA), New York (NY)

On-site
USD 166,000 - 207,000
Health insurance
Dental coverage
Vision coverage
+2
LLM Benchmarking Research Scientist
LLM Benchmarking Research Scientist

Anyone AI Inc. • Northern (KY)

Hybrid
USD 110,000 - 160,000
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
Research Scientist
Research Scientist

Anyone AI Inc. • Northern (KY)

On-site
USD 110,000 - 160,000
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI • United States

On-site
USD 181,000 - 226,000