Research Engineer, Code RL (Reinforcement Learning)

Anthropic

New York (NY)

Hybrid

USD 500,000 - 850,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is hiring a Research Engineer for the Code RL team in New York. Applicants should possess strong software engineering skills and deep expertise in Python to design RL environments and run training experiments. The position requires balancing research and engineering responsibilities while ensuring code quality and performance.

With an annual salary ranging from $500,000 to $850,000, Anthropic aims to build reliable and beneficial AI systems. Candidates must have at least a Bachelor's degree or equivalent experience.

Qualifications

  • Have strong software engineering skills and deep Python expertise, including async and concurrent programming.
  • Balance research exploration with engineering implementation, and rigorously design experiments and interpret results.
  • Experience with reinforcement learning, RLHF, or large-scale distributed training.

Responsibilities

  • Design RL environments and coding tasks, build reward signals and verifiers.
  • Run training experiments on frontier models and diagnose model behavior.
  • Enhance speed and reliability of training pipelines.

Skills

Strong software engineering skills
Deep Python expertise
Research exploration and engineering implementation
Code quality and performance awareness

Education

Bachelor’s degree or equivalent

Tools

PyTorch
CUDA / GPU or TPU

Job description

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society.

About the RL Teams

The Reinforcement Learning teams are critical to advancing our AI systems, contributing to all Claude models and impacting autonomy and coding capabilities. Core work includes:

  • Developing systems that enable models to use computers effectively
  • Advancing code generation through reinforcement learning
  • Pioneering fundamental RL research for large language models
  • Building scalable RL infrastructure and training methodologies
  • Enhancing model reasoning capabilities
About The Role

We are hiring a Research Engineer for the Code RL team. You will design RL environments and coding tasks, build reward signals and verifiers that capture the essence of "good code", run training experiments on frontier models, diagnose model behavior, and improve speed and reliability of training pipelines. The role combines research and engineering across multiple focus areas, including agentic coding behaviors, long‑horizon autonomous engineering, and high‑performance code for accelerators.

You May Be a Good Fit If You
  • Have strong software‑engineering skills and deep Python expertise, including async and concurrent programming.
  • Own systems end to end and debug across the stack.
  • Balance research exploration with engineering implementation, and rigorously design experiments and interpret results.
  • Care about code quality, testing, and performance.
  • Are passionate about AI’s impact and committed to building safe and beneficial systems.
Strong Candidates May Also Have
  • Experience with reinforcement learning, RLHF, post‑training, or LLM finetuning.
  • Built coding agents, code‑execution sandboxes, evaluation harnesses, verifiers, or developer tooling.
  • Background in program analysis, testing, verification, compilers, or formal methods.
  • Experience with PyTorch and large‑scale distributed training; performance profiling and ML system optimization.
  • CUDA / GPU or TPU kernel experience and accelerator‑performance intuition.
  • Experience with virtualization and sandboxed code execution environments.
Related Roles
  • Research Engineer, Performance RL — teach Claude to write correct, fast code for accelerators.
  • Research Engineer, Universes — long‑horizon, ultra‑realistic agentic training environments.
  • Research Engineer, Cybersecurity RL — RL for security‑relevant coding capabilities.
Annual Salary

$500,000—$850,000 USD

Logistics
  • Minimum education: Bachelor’s degree or equivalent combination of education, training, and experience.
  • Required field of study: Relevant coursework, training, or professional experience.
  • Minimum years of experience: Varies by internal job level.
  • Location‑based hybrid policy: Staff are expected to be in an office at least 25% of the time.
  • Visa sponsorship: We sponsor visas where possible and will make every reasonable effort to obtain a visa for an offer recipient.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Code RL (Reinforcement Learning)
Research Engineer, Code RL (Reinforcement Learning)

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Research Engineer, Code RL (Reinforcement Learning) San Francisco, CA | New York City, NY
Research Engineer, Code RL (Reinforcement Learning) San Francisco, CA | New York City, NY

Anthropic • San Francisco (CA)

On-site
USD 500,000 - 850,000
Research Engineer, Code RL (Reinforcement Learning)
Research Engineer, Code RL (Reinforcement Learning)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Research Engineer, Code RL (Reinforcement Learning) Anthropic San Francisco, CA | New York City, NY
Research Engineer, Code RL (Reinforcement Learning) Anthropic San Francisco, CA | New York City, NY

Neura Market • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Equity donation matching
Generous vacation
Parental leave
+2
Research Engineer, Code RL (Reinforcement Learning)
Research Engineer, Code RL (Reinforcement Learning)

Jobzhr • San Francisco (CA), Northern (KY)

Hybrid
USD 500,000 - 850,000
Research Engineer, Performance RL (Reinforcement Learning)
Research Engineer, Performance RL (Reinforcement Learning)

Menlo Ventures • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Research Engineer, Chip Design RL (Reinforcement Learning)
Research Engineer, Chip Design RL (Reinforcement Learning)

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Flexible hours
Generous vacation
Parental leave
+2
Research Engineer, Performance RL (Reinforcement Learning)
Research Engineer, Performance RL (Reinforcement Learning)

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 850,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
+1
Software Engineer, RL Data
Software Engineer, RL Data

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Software Engineer, RL Data
Software Engineer, RL Data

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000