Evaluation Platform Engineer: Build Scalable ML Benchmarks

Thinking Machines Lab

San Francisco (CA)

On-site

USD 300,000 - 475,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
Relocation support

Job summary

Thinking Machines, based in San Francisco, seeks an experienced engineer to design and build a self-serve evaluation platform for frontier AI models. You will work across Python frameworks, data pipelines, APIs, and front-end interfaces to enable reproducible evaluations and insights that guide research and model releases.

You will collaborate with pre-training, post-training, and applied teams to improve evaluation methods and turn results into decisive research actions within a fast-paced

Qualifications

  • A bachelor's degree, or equivalent practical experience, in CS, engineering, ML, or a related field.
  • Two years of post-grad work as a software or ML engineer.
  • Experience building evaluations, benchmarks, graders, or model-quality systems.
  • Strong software engineering fundamentals for reliable, maintainable systems.
  • Proficiency in Python or Rust; frontend uses React/TypeScript.
  • Experience with databases, data pipelines, or distributed data infrastructures.
  • Ability to own projects from discovery to deployment.
  • Experience collaborating with researchers and cross-functional partners.

Responsibilities

  • Design, build, and maintain evaluation platform end to end.
  • Work across evaluation libraries, data pipelines, APIs, and front-end apps.
  • Create abstractions for evaluation tasks, environments, datasets, and outputs.
  • Ensure reproducibility with versioning, provenance and observability.
  • Collaborate with researchers to streamline bespoke workflows.
  • Work with internal research platform team to integrate tools.

Skills

Python
Rust
React/TypeScript
Data pipelines
Distributed systems
Backend development
Cross-functional teamwork

Education

Bachelor's degree in CS or related

Tools

PostgreSQL
Airflow
Docker
Kubernetes

Job description

Thinking Machines, based in San Francisco, seeks an experienced engineer to design and build a self-serve evaluation platform for frontier AI models. You will work across Python frameworks, data pipelines, APIs, and front-end interfaces to enable reproducible evaluations and insights that guide research and model releases.

You will collaborate with pre-training, post-training, and applied teams to improve evaluation methods and turn results into decisive research actions within a fast-paced

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Tools Platform Engineer (ML Infra)
Research Tools Platform Engineer (ML Infra)

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
ML Evaluation Engineer: Benchmark & Model Quality
ML Evaluation Engineer: Benchmark & Model Quality

Reducto • San Francisco (CA)

On-site
USD 100,000 - 130,000
Unlimited PTO
Daily free lunch
Reimbursed transportation
+3
Software Engineer, Evaluation Platform / Infra
Software Engineer, Evaluation Platform / Infra

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Forward Deployed ML Engineer: Benchmarks & Evaluations
Forward Deployed ML Engineer: Benchmarks & Evaluations

Protege • United States

Remote
USD 120,000 - 180,000
Staff Engineer - AI Evaluation & Metrics Platform
Staff Engineer - AI Evaluation & Metrics Platform

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 200,000
End-to-End Platform Engineer for AI Benchmarking
End-to-End Platform Engineer for AI Benchmarking

David Joseph & Company • San Francisco (CA)

On-site
USD 150,000 - 210,000
Platform Engineer for Scalable ML Evaluations (Remote)
Platform Engineer for Scalable ML Evaluations (Remote)

Jaide Health • United States

On-site
USD 100,000 - 150,000
Fully remote work & flexible hours
37 days/year of vacation & holidays
Health insurance allowance
+4
ML Platform Engineer: Scalable AI Infrastructure
ML Platform Engineer: Scalable AI Infrastructure

Bjak • Germany (OH)

On-site
USD 120,000 - 180,000
AI Benchmarking Engineer — Evaluation & Failure Analysis
AI Benchmarking Engineer — Evaluation & Failure Analysis

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+6
Senior ML Platform Engineer - Scale Experiments & AI
Senior ML Platform Engineer - Scale Experiments & AI

EngineersOfAI • United States

Hybrid
USD 100,000 - 130,000