AI Experimentation Engineer for Agent Decision Systems

Radical AI

New York (NY)

On-site

USD 140,000 - 210,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Radical AI is seeking an engineering-focused researcher to design evaluation datasets, replay historical runs, and ship improvements to our agent's planning and control loops. You will own evaluation and experimentation systems and collaborate with platform engineers.

The profile is engineering-first, ML-strong, and research-capable. You should design data schemas and CI pipelines, and turn ML ideas into controlled experiments that ship improvements with production-quality results.

Qualifications

  • 4–8 years building production ML systems, research infrastructure, or data-intensive backends.
  • Excellent Python and strong software engineering discipline: testing, typing, packaging, code review, CI
  • Direct experience with ML evaluation and experimentation: offline evaluation harnesses, dataset and split design, metric design, and sound reasoning about noisy or underpowered results
  • Data pipeline skills: schema design, versioning, orchestration (Airflow, Prefect, Dagster, Ray, or equivalent), object storage
  • A reproducibility instinct: pinned environments, seeded runs, versioned artifacts, results that hold up when someone else reruns them six months later
  • The ability to take ambiguous system behavior, reduce it to a controlled experiment, and write up what the result does and does not support
  • Ability to turn ideas into working prototypes quickly, then harden successful ones into reliable production systems
  • Initiative and ownership in an environment where the roadmap is still being written

Responsibilities

  • Formulate hypotheses about agent behavior and design controlled experiments, implementing and evaluating improvements to planning, tool use, and decision-making.
  • Reproduce historical agent runs from recorded context for fair model comparisons under consistent conditions.
  • Develop versioned evaluation datasets, curated task sets, splits, and ground truth that evolve without invalidating past results.
  • Build evaluation harness and metrics: decision quality, task success, recovery from error, cost and latency.
  • Instrument agent runs, characterize failure modes, and convert recurring failures into regression coverage.
  • Prototype, test, and ship control policies for when the agent revises an approach or escalates to a scientist.
  • Create experiment infrastructure: data pipelines, artifact and dataset versioning, experiment tracking to enable quick clean comparisons.
  • Collaborate with scientists and platform engineers to connect agent behavior to Workcell outcomes.

Skills

Python programming
Production ML systems
CI/testing
Experiment design
Data pipelines
Versioned experiments
LLM agents concepts

Tools

Airflow
Prefect
Dagster
Ray
Langfuse
MLflow
Weigths & Biases
DVC
OpenTelemetry
Prometheus
Grafana
Datadog

Job description

Radical AI is seeking an engineering-focused researcher to design evaluation datasets, replay historical runs, and ship improvements to our agent's planning and control loops. You will own evaluation and experimentation systems and collaborate with platform engineers.

The profile is engineering-first, ML-strong, and research-capable. You should design data schemas and CI pipelines, and turn ML ideas into controlled experiments that ship improvements with production-quality results.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Research Engineer
ML Research Engineer

Radical AI • New York (NY)

On-site
USD 140,000 - 210,000
ML Research Engineer Brooklyn, NY On-site $235,000 – $295,000 · per year · USD
ML Research Engineer Brooklyn, NY On-site $235,000 – $295,000 · per year · USD

Nodi • New York (NY)

Hybrid
USD 235,000 - 295,000
Applied AI Engineer
Applied AI Engineer

Judgment Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Staff Research Engineer: Scalable Multi-Agent Systems
Staff Research Engineer: Scalable Multi-Agent Systems

Anthropic • New York (NY)

On-site
USD 500,000 - 850,000
Remote AI Agent Evaluation Engineer
Remote AI Agent Evaluation Engineer

EPAM Systems Inc • United States

Remote
USD 140,000 - 190,000
Research Engineer, RL Environments and Infrastructure
Research Engineer, RL Environments and Infrastructure

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
AI Evaluation Engineer: RL Environments & Agents
AI Evaluation Engineer: RL Environments & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior AI Engineer: ML Systems & Innovation Lead
Senior AI Engineer: ML Systems & Innovation Lead

redalpha • Maryland

On-site
USD 170,000 - 230,000
401k matching
Paid time off
Health, dental, vision
+3
RL Systems Engineer: Environments & High-Throughput Pipelines
RL Systems Engineer: Environments & High-Throughput Pipelines

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
AI Research Engineer Agentic AI
AI Research Engineer Agentic AI

Robert Bosch Group • Sunnyvale (CA)

On-site
USD 165,000 - 180,000
Premium health coverage
401(k) with generous matching
Ample paid time off
+2