RL & Agentic ML Researcher: Data-Centric Evaluation

Protege

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Protege is seeking a Machine Learning Researcher focused on RL and agentic systems to design and evaluate datasets, tasks, and environments for high-quality AI training data. You will work with research and engineering teams to translate real-world workflows into datasets, benchmarks, and evaluation assets that illuminate model behavior in realistic settings.

You will develop frameworks to measure dataset quality, design robust benchmarks, and validate tooling for scalable evaluation.

Qualifications

  • PhD or equivalent Master’s degree plus 4+ years of industry experience in ML, CS, statistics, or related fields.
  • Strong understanding of AI model training pipelines and data’s role in performance.
  • Experience with large, unstructured datasets used for ML.
  • Experience with reinforcement learning, agentic systems, or multi-step evaluation.
  • Experience designing tasks, benchmarks, or evaluation environments for real-world model behavior.
  • Strong intuition for realism, coverage, fidelity, and meaningful outcome structure in datasets.
  • Strong experimental design, evaluation benchmarking, and data-validation skills.
  • High ownership and ability to independently solve high-impact problems.

Responsibilities

  • Design and build datasets, tasks, and environments for benchmarking agentic systems.
  • Translate real-world workflows into structured tasks and verifiable outcomes for evaluation of AI systems.
  • Develop frameworks to assess diversity, realism, coverage, fidelity, and usefulness of datasets.
  • Benchmark model behavior in RL and agentic settings, linking failures to data/design gaps.
  • Build scalable evaluation tooling and improve reproducible experimentation infrastructure.
  • Collaborate with research, engineering, and product to improve evaluation methodology and data quality.
  • Represent DataLab’s perspective in cross-functional discussions on dataset quality and benchmark design.

Skills

ML research
Reinforcement learning
Agentic systems
Experiment design
Data evaluation
Task benchmarking

Education

PhD or Master’s + 4+ years

Tools

Harbor
RL frameworks
Trajectory data tools

Job description

Protege is seeking a Machine Learning Researcher focused on RL and agentic systems to design and evaluate datasets, tasks, and environments for high-quality AI training data. You will work with research and engineering teams to translate real-world workflows into datasets, benchmarks, and evaluation assets that illuminate model behavior in realistic settings.

You will develop frameworks to measure dataset quality, design robust benchmarks, and validate tooling for scalable evaluation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

RL & Agentic AI Researcher — Benchmarking Data Quality
RL & Agentic AI Researcher — Benchmarking Data Quality

Protege • United States

On-site
USD 110,000 - 140,000
RL & Data Benchmarking Researcher
RL & Data Benchmarking Researcher

protege • United States

Remote
USD 150,000 - 230,000
ML Researcher: RL & Agentic Systems Benchmarking
ML Researcher: RL & Agentic Systems Benchmarking

Protege • New City (NY)

On-site
USD 120,000 - 160,000
Machine Learning Researcher, RL & Agentic Systems
Machine Learning Researcher, RL & Agentic Systems

Protege • United States

On-site
USD 110,000 - 140,000
Machine Learning Researcher, RL and Agentic
Machine Learning Researcher, RL and Agentic

Protege • United States

Remote
USD 140,000 - 190,000
Remote AI Research Engineer — RL & LLM Systems
Remote AI Research Engineer — RL & LLM Systems

Not specified • United States

Remote
USD 120,000 - 180,000
AI Research Engineer: Post-Training & Agentic RL
AI Research Engineer: Post-Training & Agentic RL

Oho Group • San Francisco (CA)

On-site
USD 160,000 - 230,000
Remote RL Research Intern — Agentic AI & LLMs
Remote RL Research Intern — Agentic AI & LLMs

Centific Global Solutions, Inc. • United States

On-site
USD 48,216 - 61,992
Competitive stipend
Mentorship from researchers
Access to modern GPU infrastructure
Remote AI Research Engineer - RL & Model Evaluation
Remote AI Research Engineer - RL & Model Evaluation

Appen Limited • United States

Remote
USD 120,000 - 180,000
Machine Learning Researcher - RL and Agentic
Machine Learning Researcher - RL and Agentic

Protege • New City (NY)

On-site
USD 120,000 - 160,000