ML Research Engineer

Radical AI

New York (NY)

On-site

USD 140,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Radical AI is seeking an engineering-focused researcher to design evaluation datasets, replay historical runs, and ship improvements to our agent's planning and control loops. You will own evaluation and experimentation systems and collaborate with platform engineers.

The profile is engineering-first, ML-strong, and research-capable. You should design data schemas and CI pipelines, and turn ML ideas into controlled experiments that ship improvements with production-quality results.

Qualifications

  • 4–8 years building production ML systems, research infrastructure, or data-intensive backends.
  • Excellent Python and strong software engineering discipline: testing, typing, packaging, code review, CI
  • Direct experience with ML evaluation and experimentation: offline evaluation harnesses, dataset and split design, metric design, and sound reasoning about noisy or underpowered results
  • Data pipeline skills: schema design, versioning, orchestration (Airflow, Prefect, Dagster, Ray, or equivalent), object storage
  • A reproducibility instinct: pinned environments, seeded runs, versioned artifacts, results that hold up when someone else reruns them six months later
  • The ability to take ambiguous system behavior, reduce it to a controlled experiment, and write up what the result does and does not support
  • Ability to turn ideas into working prototypes quickly, then harden successful ones into reliable production systems
  • Initiative and ownership in an environment where the roadmap is still being written

Responsibilities

  • Formulate hypotheses about agent behavior and design controlled experiments, implementing and evaluating improvements to planning, tool use, and decision-making.
  • Reproduce historical agent runs from recorded context for fair model comparisons under consistent conditions.
  • Develop versioned evaluation datasets, curated task sets, splits, and ground truth that evolve without invalidating past results.
  • Build evaluation harness and metrics: decision quality, task success, recovery from error, cost and latency.
  • Instrument agent runs, characterize failure modes, and convert recurring failures into regression coverage.
  • Prototype, test, and ship control policies for when the agent revises an approach or escalates to a scientist.
  • Create experiment infrastructure: data pipelines, artifact and dataset versioning, experiment tracking to enable quick clean comparisons.
  • Collaborate with scientists and platform engineers to connect agent behavior to Workcell outcomes.

Skills

Python programming
Production ML systems
CI/testing
Experiment design
Data pipelines
Versioned experiments
LLM agents concepts

Tools

Airflow
Prefect
Dagster
Ray
Langfuse
MLflow
Weigths & Biases
DVC
OpenTelemetry
Prometheus
Grafana
Datadog

Job description

Our scientific agent decides what experiment to run next. This role builds the machinery that lets us understand, evaluate, and continuously improve the agent's decisions.


You will work closely with the Principal Scientist to turn hypotheses about agent behavior into reproducible experiments and production improvements. Your initial focus will be evaluation datasets, replay, trajectory analysis, and the policies governing how the agent plans, responds to results, and recovers from failure.


You will own the evaluation and experimentation systems, contribute research ideas, and carry promising changes from prototype through production in partnership with platform engineers.


The profile is engineering-first, ML-strong, and research-capable. You should be as comfortable designing a data schema and CI pipeline as turning an ML or scientific idea into a controlled experiment and shipping the resulting improvement.


What You'll Work On


  • Research and experimentation: work with the Principal Scientist to formulate hypotheses about agent behavior, design controlled experiments, and implement and evaluate improvements to planning, tool use, and decision-making.

  • Reproducible replay: reconstruct historical agent runs from recorded context so models, prompts, and policies can be compared under consistent conditions, with clear limits on what those comparisons establish.

  • Versioned evaluation datasets: curated task sets, splits, and ground truth that evolve without invalidating past results

  • Evaluation harness and metrics: decision quality, task success, recovery from error, cost, and latency, with honest treatment of noise and small sample sizes

  • Trajectory analysis: instrument agent runs, characterize failure modes, and convert recurring failures into regression coverage

  • Control policies: prototype, test, and ship the logic that decides when the agent revises an approach, pivots, escalates to a scientist, or terminates

  • Experiment infrastructure: data pipelines, artifact and dataset versioning, experiment tracking, and the tooling that lets the team run a clean comparison in minutes rather than days

  • Close partnership with our scientists and platform engineers to connect agent behavior to what actually happens in the Workcell


What We're Looking For

Required:



  • 4-8 years building production ML systems, research infrastructure, or data-intensive backends; strong enough to design, build, and ship end-to-end

  • Excellent Python and strong software engineering discipline: testing, typing, packaging, code review, CI

  • Direct experience with ML evaluation and experimentation: offline evaluation harnesses, dataset and split design, metric design, and sound reasoning about noisy or underpowered results

  • Data pipeline skills: schema design, versioning, orchestration (Airflow, Prefect, Dagster, Ray, or equivalent), object storage

  • A reproducibility instinct: pinned environments, seeded runs, versioned artifacts, results that hold up when someone else reruns them six months later

  • The ability to take ambiguous system behavior, reduce it to a controlled experiment, and write up what the result does and does not support

  • Ability to turn ideas into working prototypes quickly, then harden successful ones into reliable production systems

  • Initiative and ownership in an environment where the roadmap is still being written


Nice to have:


  • Experience with LLM agents: tool calling, planning loops, context management, trajectory logging, and agent evaluation

  • Background in sequential decision making: bandits, Bayesian optimization, reinforcement learning, or off-policy evaluation

  • Experiment tracking, agent observability, and artifact versioning tools (Langfuse, MLflow, Weights and Biases, DVC)

  • Scientific computing, or experience with materials, chemistry, or instrument data

  • Experience operating large-scale inference systems, including GPU scheduling and distributed execution

  • Observability experience (OpenTelemetry, Prometheus, Grafana, Datadog)

  • Prior work in lab automation, robotics, or a self-driving lab


What Success Looks Like

In your first six to twelve months:



  • Representative historical agent runs can be replayed against recorded context and compared with candidate models, prompts, or policies, with clear limits on those comparisons.

  • A versioned evaluation suite gates changes to the agent, and the team trusts it enough to act on the result

  • Recurring failure modes are documented, categorized, and covered by regression tests.

  • Scientific and agent hypotheses can be turned into controlled experiments quickly, with clear conclusions about what the evidence supports and what to try next.

  • At least one meaningful agent or control improvement is shipped to production, with measured gains in decision quality or scientific throughput.


Radical AI is an equal opportunity employer. We do not discriminate on the basis of race, color, ancestry, national origin, religion, sex, age, sexual orientation, gender identity and expression, marital status, disability, or veteran status.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Research Engineer
ML Research Engineer

Nodi • New York (NY)

On-site
USD 235,000 - 295,000
Member of Technical Staff, Post-Training
Member of Technical Staff, Post-Training

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Meals and office benefits
Visa sponsorship
ML/AI Research Engineer — Agentic AI Lab (Founding Team)
ML/AI Research Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Staff Software Engineer, Agent Eval Platform
Staff Software Engineer, Agent Eval Platform

Servicenow • Santa Clara (CA)

On-site
USD 180,000 - 320,000
AI Agents Applied Research/Engineering Lead - Vice President
AI Agents Applied Research/Engineering Lead - Vice President

JPMorgan Chase & Co. • New York (NY)

On-site
USD 180,000 - 240,000
MTS - Engineering
MTS - Engineering

Collinear AI • Sunnyvale (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 190,000 - 230,000
MACHINE LEARNING ENGINEER (GENERAL)
MACHINE LEARNING ENGINEER (GENERAL)

MakerMaker.AI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 270,000
MACHINE LEARNING ENGINEER (GENERAL)
MACHINE LEARNING ENGINEER (GENERAL)

MakerMaker • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity