Frontier AI Benchmarks Engineer — Field‑Facing Data & Eval

Alexander Chapman

United States

On-site

USD 140,000 - 190,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Alexander Chapman is seeking a Forward Deployed ML Engineer, Benchmarks & Evaluations to help build the infrastructure behind AI evaluations and data benchmarks. You will be one of the first engineers in this high-ownership, fast-moving team, collaborating with the GM, researchers and customers to deploy repeatable evaluation products.

The role focuses on building LLM benchmarks, data pipelines, sandboxed evaluation environments, and close collaboration with customers to shape evaluation

Qualifications

  • 4+ years of engineering experience.
  • Comfortable in high-ambiguity, fast-moving environments.
  • Strong communication and ability to work directly with customers.

Responsibilities

  • Build and improve LLM/AI benchmarks and evaluation systems.
  • Develop backend infrastructure for data pipelines, orchestration, storage, and execution.
  • Build sandboxed environments for agentic evaluations, including tool use and code execution.
  • Work directly with customers to identify evaluation needs and turn them into repeatable products.
  • Partner with researchers on new datasets, benchmarks, and evaluation methodologies.

Skills

4+ years engineering experience

Job description

Alexander Chapman is seeking a Forward Deployed ML Engineer, Benchmarks & Evaluations to help build the infrastructure behind AI evaluations and data benchmarks. You will be one of the first engineers in this high-ownership, fast-moving team, collaborating with the GM, researchers and customers to deploy repeatable evaluation products.

The role focuses on building LLM benchmarks, data pipelines, sandboxed evaluation environments, and close collaboration with customers to shape evaluation

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Forward Deployed ML Engineer
Forward Deployed ML Engineer

Alexander Chapman • United States

On-site
USD 140,000 - 190,000
Remote Forward-Deployed ML Engineer — Benchmark & Eval
Remote Forward-Deployed ML Engineer — Benchmark & Eval

Advatix • Northern (KY)

Hybrid
USD 170,000 - 270,000
Frontier AI Research Scientist: Benchmarks & Data
Frontier AI Research Scientist: Benchmarks & Data

turing • San Francisco (CA)

On-site
USD 250,000 - 350,000
Remote ML Engineer: Build Benchmarks & Production Pipelines
Remote ML Engineer: Build Benchmarks & Production Pipelines

raydar • Northern (KY)

Hybrid
USD 170,000 - 270,000
Competitive equity
Forward Deployed ML Engineer: Benchmarks & Evaluations
Forward Deployed ML Engineer: Benchmarks & Evaluations

Protege • United States

Remote
USD 120,000 - 180,000
Frontier AI ML Engineer — Evaluation & Production Systems
Frontier AI ML Engineer — Evaluation & Production Systems

Obsidian • New York (NY)

Remote
USD 4,000 - 7,000
Benchmark Research Engineer for Frontier AI
Benchmark Research Engineer for Frontier AI

HUD • San Francisco (CA)

On-site
USD 140,000 - 190,000
Competitive compensation
Top-tier medical, dental, and vision (
Lunch and dinner in office
+5
Evaluation Platform Engineer: Build Scalable ML Benchmarks
Evaluation Platform Engineer: Build Scalable ML Benchmarks

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Infra Engineer for Frontier AI Biology Benchmarks
Infra Engineer for Frontier AI Biology Benchmarks

LatchBio • San Francisco (CA)

On-site
USD 180,000 - 250,000
Unlimited PTO
Premium health plan
Office in San Francisco
+3
Member of Technical Staff - ML Evaluation & Benchmarking
Member of Technical Staff - ML Evaluation & Benchmarking

Arena Physica • New York (NY)

On-site
USD 175,000 - 250,000
Premium health benefits
401(k)
Unlimited PTO
+2