Research Engineer, RL Scaling Science

Humanloop

Greater London

Hybrid

GBP 57,000 - 73,000

Full time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity donation matching
Generous vacation and parental leave
Flexible working hours
Hybrid work policy
Visa sponsorship where possible

Job summary

Anthropic is seeking a research-focused engineer to design, run, and interpret large-scale reinforcement learning experiments. You will work at the intersection of research and systems, debugging complex issues and scaling experiments to production-ready pipelines.

The role emphasizes strong Python skills, experience with distributed ML systems, and a commitment to responsible AI. A hybrid working arrangement is expected, with visa sponsorship where possible.

Qualifications

  • Strong empirical research skills in RL, or closely related ML areas.
  • Proven ability to own large experiments end-to-end from design to interpretation.
  • Proficiency in Python and experience with large-scale or distributed ML systems.
  • Comfort operating at the research/systems boundary and debugging at scale.
  • Concern for societal impacts of AI and responsible scaling.
  • Field relevant to the role demonstrated through coursework or experience.

Responsibilities

  • Design, run, and interpret large-scale RL experiments; reason about data validity.
  • Investigate RL progress as horizon, compute, and model size grow.
  • Build and maintain benchmarks for long-horizon RL to measure progress.
  • Translate findings into production training recipes with robustness checks.
  • Debug complex issues at the research–infrastructure interface, at scale.
  • Collaborate with adjacent RL teams to advance the RL stack.

Skills

Reinforcement Learning
Large-scale ML training
Experiment ownership
Python proficiency
Distributed ML systems
Research–systems boundary
Societal impact awareness

Education

Bachelor's degree or equivalent

Tools

Python

Job description

Salary: £57,000 - 73,000 per year

Requirements
  • We require strong empirical research skills in Reinforcement Learning, large-scale ML training, or a closely adjacent area.
  • We require demonstrated ability to own large experiments end-to-end, from design through interpretation.
  • We require proficiency in Python and experience working with large-scale or distributed ML systems.
  • We require comfort operating at the research/systems boundary, including debugging where the two meet.
  • We require concern for the societal impacts of AI and responsible scaling.
  • We require a bachelors degree or an equivalent combination of education, training, and/or experience.
  • We require a field relevant to the role, as demonstrated through coursework, training, or professional experience.
  • We require years of experience aligned with the internal job level requirements for the position.
Responsibilities
  • We design, run, and interpret large-scale RL experiments, reasoning rigorously about what the data does and does not show.
  • We investigate how RL improves as horizon, compute, and model size grow.
  • We build and maintain benchmarks for long-horizon RL so progress is measurable and reproducible.
  • We translate validated findings into production training recipes, exercising judgment about when a result is robust enough to ship.
  • We debug complex issues at the seam where research meets infrastructure, including failures that only appear at scale.
  • We partner closely with adjacent RL teams across research and engineering and advance our overall RL stack.
Technologies
  • AI
  • Python
More

We are Anthropic, a public benefit corporation headquartered in San Francisco. Our mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for our users and for society. Our team is a quickly growing group of researchers, engineers, policy experts, and business leaders working together on big-science AI research and long-term goals such as steerable, trustworthy AI. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a collaborative office environment. We operate with a hybrid policy and expect staff to be in one of our offices at least 25% of the time, and we sponsor visas where possible.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, RL Scaling Science
Research Engineer, RL Scaling Science

Anthropic • Greater London

Hybrid
GBP 375,000 - 640,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
+1
Research Engineer, RL Scaling Science New London, UK
Research Engineer, RL Scaling Science New London, UK

Alcides Fonseca • Greater London

Hybrid
GBP 70,000 - 90,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Research Engineer, Machine Learning (RL Velocity)
Research Engineer, Machine Learning (RL Velocity)

Humanloop • Greater London

Hybrid
GBP 57,000 - 73,000
Visa sponsorship
Flexible working hours
Hybrid work policy
+1
RL Scaling Engineer — Benchmarks, Production Recipes
RL Scaling Engineer — Benchmarks, Production Recipes

Anthropic • Greater London

Hybrid
GBP 375,000 - 640,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
+1
Research Engineer, Machine Learning (Reinforcement Learning)
Research Engineer, Machine Learning (Reinforcement Learning)

Anthropic • Greater London

Hybrid
GBP 260,000 - 630,000
Research Engineer, Machine Learning (RL Velocity)
Research Engineer, Machine Learning (RL Velocity)

Menlo Ventures • Greater London

On-site
GBP 370,000 - 630,000
RL Research Engineer: Safe, Scalable AI Systems
RL Research Engineer: Safe, Scalable AI Systems

Anthropic • Greater London

Hybrid
GBP 260,000 - 630,000
Frontier RL Research Engineer
Frontier RL Research Engineer

Alcides Fonseca • Greater London

Hybrid
GBP 70,000 - 90,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
RL Scaling Scientist | End-to-End Research to Production
RL Scaling Scientist | End-to-End Research to Production

Humanloop • Greater London

Hybrid
GBP 57,000 - 73,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+2
Research Engineer
Research Engineer

Adecco • Greater London

On-site
GBP 60,000 - 82,000