Research Scientist, Agentic RL & Scalable AI

Goaly

Menlo Park, Northern (CA, KY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Meals and office benefits

Job summary

Goaly is seeking researchers at the intersection of agentic reinforcement learning and production systems. You will identify high-leverage questions, design decisive experiments, and translate results into model improvements and reusable systems.

This role suits candidates completing a PhD or with equivalent original research; you will work closely with Post-Training, RL Systems, Training, Backend & Product engineers to scale ideas within production constraints.

Qualifications

  • PhD or equivalent record of original, rigorous research.
  • Strong research record in ML/RL/LLMs, agents, or ML systems.
  • Excellent Python skills with a modern DL framework (PyTorch/JAX).
  • Experimental rigor: define falsifiable questions, measure, control confounders, interpret results.

Responsibilities

  • Formulate high-leverage research questions about agentic capability, RL, reward design, and scaling.
  • Design experiments, ablations, controls, and evaluations that separate true improvement from noise.
  • Implement methods in modern DL frameworks and integrate with training/evaluation systems.
  • Build or improve datasets, environments, verifiers, and evaluations for coding, tool use, reasoning, and interaction.
  • Analyze trajectories and model behavior; develop failure taxonomies and testable hypotheses.
  • Collaborate with Post-Training, RL Systems, Training, and Backend & Product engineers to scale ideas.
  • Communicate findings in internal docs and reviews; contribute to papers and open-source releases.
  • Help shape the research roadmap by identifying reusable evaluation assets and critical uncertainties.

Skills

PhD in CS/ML/Math
Research track record
Python
Experimental rigor
Engineering ability
Communication

Education

PhD in Computer Science, ML, statistics, mathematics

Tools

PyTorch
JAX

Job description

Goaly is seeking researchers at the intersection of agentic reinforcement learning and production systems. You will identify high-leverage questions, design decisive experiments, and translate results into model improvements and reusable systems.

This role suits candidates completing a PhD or with equivalent original research; you will work closely with Post-Training, RL Systems, Training, Backend & Product engineers to scale ideas within production constraints.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff AI Research Engineer - Post-Training Agents
Staff AI Research Engineer - Post-Training Agents

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Meals and office benefits
Visa sponsorship
AI Research Engineer: Post-Training & Agentic RL
AI Research Engineer: Post-Training & Agentic RL

Oho Group • San Francisco (CA)

On-site
USD 160,000 - 230,000
Pioneering AI Infrastructure Engineer
Pioneering AI Infrastructure Engineer

Goaly • Palo Alto (CA)

On-site
USD 140,000 - 210,000
Principal Research Scientist: AI Systems & RL in Production
Principal Research Scientist: AI Systems & RL in Production

Centific • East Palo Alto (CA)

On-site
USD 250,000 - 300,000
Lead RL & Agentic AI Research—On-Device & Tools
Lead RL & Agentic AI Research—On-Device & Tools

Socket.dev • Cupertino (CA)

On-site
USD 220,000 - 320,000
RL Research Engineer - Scalable, Safe AI Systems
RL Research Engineer - Scalable, Safe AI Systems

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
Remote RL Research Intern — Agentic AI & LLMs
Remote RL Research Intern — Agentic AI & LLMs

Centific Global Solutions, Inc. • United States

On-site
Competitive stipend
Mentorship from researchers
Access to modern GPU infrastructure
AIML - Machine Learning Research Lead, RL Agents, MLR
AIML - Machine Learning Research Lead, RL Agents, MLR

Socket.dev • Cupertino (CA)

On-site
USD 220,000 - 320,000
research scientist - RL
research scientist - RL

Cerebro • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lead RL Researcher for Agentic AI & Environments
Lead RL Researcher for Agentic AI & Environments

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 216,000 - 394,000
Apple stock programs
Medical coverage
Retirement benefits
+1