Agentic Data Scientist & Benchmarking Researcher

Ifm Us

Sunnyvale (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Institute of Foundation Models (IFM) is seeking a research scientist focused on data and evaluation depth for agentic models. You will own end‑to‑end data pipelines, sourcing high‑quality trajectories and tool‑use data, while building a rigorous evaluation suite to measure progress with reproducible metrics.

You will work with researchers and engineers to ensure data quality, design robust RL environments, and contribute to publications and benchmarks, advancing the state of the art in agentic

Qualifications

  • BS, MS or PhD (or equivalent experience) in Computer Science, Machine Learning or a related field.
  • 2+ years of experience with a clear emphasis on evaluations and/or training data curation for ML systems.
  • Strong Python and PyTorch development experience.
  • Demonstrated experience designing and deep‑diving into evaluations, or curating and generating training datasets — ideally both.
  • Hands‑on experience using LLM agents in your personal or professional work.
  • A habit of reading through raw data and trajectories to understand them and spot issues, and an instinct to distrust a metric until it's validated.

Responsibilities

  • Design and run evaluations of agentic capabilities — multi‑step reasoning, tool use, long‑horizon planning, computer use and safety properties — turning ambiguous notions of 'intelligence' into defensible, reproducible metrics.
  • Build and harden evaluation harnesses so benchmarks run reliably at scale against training checkpoints, with clear signal on regressions and model health.
  • Run experiments characterizing how prompting, sampling, scaffolding and environment design affect agentic performance on internal and public benchmarks.
  • Diagnose anomalous eval results mid‑training run — determine whether the cause is the model, the data, the harness or the infrastructure — and communicate the answer clearly.

Skills

Evaluation design
Data curation
LLM agents
Reading raw data
Metric validation

Education

BS/MS/PhD in CS/ML

Tools

Python
PyTorch
RL environments

Job description

Institute of Foundation Models (IFM) is seeking a research scientist focused on data and evaluation depth for agentic models. You will own end‑to‑end data pipelines, sourcing high‑quality trajectories and tool‑use data, while building a rigorous evaluation suite to measure progress with reproducible metrics.

You will work with researchers and engineers to ensure data quality, design robust RL environments, and contribute to publications and benchmarks, advancing the state of the art in agentic

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist, Agentic Data & Benchmarking
Research Scientist, Agentic Data & Benchmarking

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 180,000 - 250,000
RL & Agentic AI Researcher — Benchmarking Data Quality
RL & Agentic AI Researcher — Benchmarking Data Quality

Protege • United States

On-site
USD 110,000 - 140,000
ML Researcher: RL & Agentic Systems Benchmarking
ML Researcher: RL & Agentic Systems Benchmarking

Protege • New City (NY)

On-site
USD 120,000 - 160,000
Data Scientist - Agentic AI Systems & Fast-Paced Research
Data Scientist - Agentic AI Systems & Fast-Paced Research

Industrial and Financial Systems • Palo Alto (CA)

On-site
USD 140,000 - 150,000
Salary Range $140k–$150k annually + 0–
401(k)
Flexible PTO
Impactful Research Scientist: Data & Evaluation
Impactful Research Scientist: Data & Evaluation

bareinsights • Los Angeles (CA)

On-site
USD 110,000 - 150,000
Research Scientist, Data
Research Scientist, Data

Prophet Town • San Francisco (CA)

On-site
USD 120,000 - 160,000
Agentic AI Data Scientist — Fast Prototyping & Real-World Labs
Agentic AI Data Scientist — Fast Prototyping & Real-World Labs

IFS • Palo Alto (CA)

On-site
USD 140,000 - 150,000
Flexible paid time off, including sick
401(k) with company contribution.
Flexible spending accounts.
+3
Research Scientist (Remote/US/LATAM)
Research Scientist (Remote/US/LATAM)

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
Founding AI Research Lead - Agentic AI Lab
Founding AI Research Lead - Agentic AI Lab

Fabrion • San Francisco (CA)

On-site
USD 180,000 - 240,000
Remote AI Benchmark & Datasets Engineer
Remote AI Benchmark & Datasets Engineer

Pathway Genomics Corporation • Palo Alto (CA)

Remote
USD 150,000 - 210,000
Intellectually stimulating work environment
Work with a pioneering AI startup
Flexible remote work options