Research Scientist, Agentic Data & Benchmarking

Institute of Foundation Models

Sunnyvale (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Institute of Foundation Models is seeking a research scientist focused on data and measurement. You will own agentic data pipelines, curate trajectories and RL environments, and build evaluation suites to rigorously assess model capabilities.

You will read datasets line by line, validate metrics, and push for robust, reproducible benchmarks. Collaboration with researchers and engineers will translate capability goals into measurable data artifacts.

Qualifications

  • BS/MS/PhD in Computer Science, ML, or related field.
  • 2+ years evaluating ML systems, data curation for ML training, or RL.
  • Strong Python and PyTorch development experience.
  • Experience designing and deep-diving into evaluations or curating datasets.
  • Hands-on use of LLM agents in work or personal projects.

Responsibilities

  • Design and run evaluations of agentic capabilities with multi-step reasoning and tool use.
  • Build evaluation harnesses and scale benchmarks across training checkpoints.
  • Source, generate, and curate high-quality agentic training data and RL environments.
  • Develop QA frameworks to catch reward hacking and data contamination.
  • Contribute to technical reports, publications, and benchmarks.

Skills

Python
PyTorch
Evaluation design
Data curation
LLM agent experience
Data analysis

Education

BS/MS/PhD in CS/ML

Job description

About the Institute of Foundation Models

The Institute of Foundation Models (IFM) is a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.

As part of our team,you'llwork at the coreofcutting-edgefoundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development.You'llhelp build groundbreaking AI systems with the potential to reshape entireindustries, andcontribute to establishing MBZUAI as a global hub for high-performance computing and deep learning.

About the role

The Agents team trains advanced agentic language models that use reasoning andtool useto complete real tasks on a computer. This is a specialist role at the center of the loop that drives those models:the data we train on and the benchmarks we measure against.

You'llown the agentic data pipeline end-to-end — sourcing and generating high-quality trajectories, tool-use data, and RL environments — and the evaluation suite that tells us, rigorously and reproducibly, what our agents canactually do. These two halves are inseparable: benchmarks expose where models fail, and targeted data closes the gap. The agents are only as good as the data they learn from and the evals that keep us honest, and this role owns both.

This is a researchscientistposition for someone who wants depth in data and measurement rather than breadth across the whole stack. You should be the kind of person who reads throughdatasets line by line, distrusts a metric until the validation, and gets satisfaction from making an eval suite that nobody questions.

Key responsibilities
Benchmarking & evaluation
  • Design and run evaluations of agentic capabilities — multi-step reasoning, tool use, long-horizon planning, computer use, and safety properties — turning ambiguous notions of "intelligence" into defensible, reproducible metrics.

  • Build and harden evaluation harnesses so benchmarks run reliably at scale against training checkpoints, with clear signal on regressions and model health.

  • Run experiments characterizing how prompting, sampling, scaffolding, and environment design affect agentic performance on internal and public benchmarks.

  • Diagnose anomalous eval results mid-training run — determine whether the cause is the model, the data, the harness, or the infrastructure — and communicate the answer clearly.

Agentic data
  • Source, generate, and curate high-quality agentic training data: trajectories, tool-use traces, and task datasets for new capabilities.

  • Design and scale RL environments and reward signals, and measure their impact on model performance.

  • Manage technical relationships with external data vendors and domain experts, evaluating data quality and iterating quickly on feedback.

  • Develop QA frameworks that catch reward hacking, label noise, and contamination, keeping data and benchmark quality high.

Across both
  • Contribute to technical reports, research publications, and open-source benchmarks and tooling.

  • Partner with research and product teams to translate capability goals into measurable data and evaluation artifacts.

Qualifications
Academic qualifications
  • BS, MS, or PhD (or equivalent experience) in Computer Science, Machine Learning, or a related field.
Minimum qualifications
  • 2+ years of experience with a clear emphasis onevaluations and/ortraining-datacurationfor ML systems (related areas: LLM training/fine-tuning, RL, or distributed ML systems).

  • Strong Python and PyTorch development experience.

  • Demonstrated experience designing and deep-diving into evaluations, or curating and generating training datasets — ideally both.

  • Hands-on experience using LLM agents in your personal or professional work.

  • A habit of reading through raw data and trajectories to understand them and spot issues, and an instinct to distrust a metric until the validation.

Preferred qualifications
  • Experience with reinforcement learning, reward design, or RL environment construction for LLMs.

  • Background in statistics and experimental design — a feel for signal-to-noise, statistical power, and contamination in evaluations.

  • Experience with large-scale dataset sourcing, curation, and processing, including working with external vendors or domain experts.

  • Strong knowledge of the literature on agent evaluation, RL, LLM reasoning, and tool use.

  • Experience building or operating data pipelines and evaluation infrastructure reliable at scale (e.g

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist - Agents
Research Scientist - Agents

Ifm Us • Sunnyvale (CA)

On-site
USD 120,000 - 190,000
Comprehensive medical, dental, and vis
Bonus
401K Plan
+4
Research Scientist (Remote/US/LATAM)
Research Scientist (Remote/US/LATAM)

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
ML/AI Research Engineer — Agentic AI Lab (Founding Team)
ML/AI Research Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 170,000 - 210,000
Competitive salary
Meaningful equity (founding tier)
Sr. Machine Learning Engineer
Sr. Machine Learning Engineer

Intel • Santa Clara (CA)

On-site
USD 180,000 - 270,000
RESEARCHER, AGENTS FOR AUTOMATED DISCOVERY
RESEARCHER, AGENTS FOR AUTOMATED DISCOVERY

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Relocation support
Health insurance
Lunch and snacks provided
+2
Member of Technical Staff [Research]
Member of Technical Staff [Research]

NeoCognition Inc. • Palo Alto (CA)

On-site
USD 120,000 - 150,000
LLM Training & Model Development Engineer
LLM Training & Model Development Engineer

InOpTra Digital • United States

Remote
USD 90,000 - 120,000
Competitive salary
Opportunity for remote work
Health benefits
ML Research Engineer - PhD - AI Trainer
ML Research Engineer - PhD - AI Trainer

Obsidian • Seattle (WA)

Hybrid
USD 100,000 - 150,000