Research Engineer, RL Scaling Science

Anthropic

Greater London

Hybrid

GBP 375,000 - 640,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Generous vacation and parental leave
Flexible working hours
Collaborative office space

Job summary

Anthropic is seeking a Research Engineer to join their RL Scaling Science team in Greater London. You will focus on designing and running large-scale experiments to enhance understanding and application of Reinforcement Learning.

The ideal candidate has strong empirical skills and proficiency in Python, alongside experience with large-scale ML systems. A Bachelor's degree or equivalent is required. The role offers a competitive salary between £375,000 and £640,000 GBP.

Qualifications

  • Strong empirical research skills in Reinforcement Learning or large-scale ML training.
  • Ability to own large experiments end-to-end.
  • Comfort operating at the research/systems boundary.

Responsibilities

  • Design, run, and interpret large-scale RL experiments.
  • Investigate how RL improves as horizon, compute, and model size grow.
  • Build and maintain benchmarks for long-horizon RL.

Skills

Empirical research skills in Reinforcement Learning
Proficiency in Python
Experience with large-scale ML systems
Debugging skills at research/systems boundary

Education

Bachelor’s degree or equivalent

Job description

About the Role

Anthropic's RL Scaling Science team studies how reinforcement learning behaves as we scale it (across model size, compute, and task horizon) and turns that understanding into the training recipes behind our frontier models. As a Research Engineer on this team, you will design and run large‑scale experiments to understand and resolve bottlenecks, build the benchmarks that make long‑horizon progress measurable, and ship validated findings directly into production training.

Key Responsibilities
  • Design, run, and interpret large‑scale RL experiments, reasoning rigorously about what the data does and doesn’t show
  • Investigate how RL improves as horizon, compute, and model size grow
  • Build and maintain benchmarks for long‑horizon RL so progress is measurable and reproducible
  • Translate validated findings into production training recipes, exercising judgment about when a result is robust enough to ship
  • Debug complex issues at the seam where research meets infrastructure – failures that only appear at scale
  • Partner closely with adjacent RL teams across research and engineering and advance our overall RL stack
Minimum Qualifications
  • Strong empirical research skills in Reinforcement Learning, large‑scale ML training, or a closely adjacent area
  • Demonstrated ability to own large experiments end‑to‑end, from design through interpretation
  • Proficiency in Python and experience working with large‑scale or distributed ML systems
  • Comfort operating at the research/systems boundary, including debugging where the two meet
  • Care about the societal impacts of AI and responsible scaling
Preferred Qualifications
  • Published or shipped work in long‑horizon RL or RL fundamentals
  • Experience translating research findings into production training recipes
  • Demonstrated large‑scale industry impact via RL interventions
  • Experience working on frontier‑scale training runs with long trajectories
Representative Projects
  • Design a benchmark suite for long‑horizon RL that distinguishes genuine capability gains from artifacts of evaluation setup
  • Take a promising experimental finding, stress‑test it across model scales, and work with training teams to land it in a production recipe
  • Investigate an unexpected scaling trend in an RL run and trace it to a root cause spanning algorithm, data, and infrastructure
Annual Salary

£375,000—£640,000 GBP

Logistics
  • Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
  • Location‑based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.
  • Visa sponsorship: We do sponsor visas; however, we may not be able to sponsor visas for every role and every candidate.
Benefits
  • Competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, RL Scaling Science New London, UK
Research Engineer, RL Scaling Science New London, UK

Alcides Fonseca • Greater London

Hybrid
GBP 70,000 - 90,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
RL Scaling Engineer — Benchmarks, Production Recipes
RL Scaling Engineer — Benchmarks, Production Recipes

Anthropic • Greater London

Hybrid
GBP 375,000 - 640,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
+1
Research Engineer, Machine Learning (RL Velocity)
Research Engineer, Machine Learning (RL Velocity)

Menlo Ventures • Greater London

On-site
GBP 370,000 - 630,000
Research Engineer, Machine Learning (Reinforcement Learning)
Research Engineer, Machine Learning (Reinforcement Learning)

Anthropic • Greater London

Hybrid
GBP 260,000 - 630,000
Member of Technical Staff - Research Scientist
Member of Technical Staff - Research Scientist

General Reasoning, Inc. • Greater London

On-site
GBP 70,000 - 90,000
Software Engineer, RL Data
Software Engineer, RL Data

United States Digital Space LLC • Greater London

Hybrid
GBP 242,000 - 368,000
Generous vacation and parental leave
Flexible working hours
Office space for collaboration
Frontier RL Research Engineer
Frontier RL Research Engineer

Alcides Fonseca • Greater London

Hybrid
GBP 70,000 - 90,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Research Engineer, Pretraining Scaling - London
Research Engineer, Pretraining Scaling - London

Anthropic • Greater London

On-site
GBP 250,000 - 435,000
Equity benefits
Visa sponsorship
Research Engineer, Pretraining Scaling - London
Research Engineer, Pretraining Scaling - London

Anthropic • Greater London

On-site
GBP 260,000 - 630,000
Research Engineer (Machine Learning, Reinforcement Learning Velocity)
Research Engineer (Machine Learning, Reinforcement Learning Velocity)

Anthropic • Greater London

On-site
GBP 120,000 - 180,000
Commuter benefits
Education stipend
Home office stipend
+2