RL Environment Data Engineer / Researcher

Eigent AI

Greater London

On-site

GBP 65,000 - 85,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Eigent AI is seeking an RL Environment Data Engineer / Researcher to design and refine reinforcement learning training environments. This role emphasizes data collection, task definition, and the implementation of anti-reward-hacking mechanisms.

The ideal candidate will have strong Python coding skills and a solid understanding of reinforcement learning. Responsibilities include collaborating with multiple teams to improve RL tasks and validation environments.

Qualifications

  • Strong coding skills in Python to build data pipelines and evaluation tools.
  • Understanding reinforcement learning and environment design.
  • Experience in data quality assessment preferred.

Responsibilities

  • Design and improve RL training environments.
  • Collect, clean, structure, and evaluate data for RL models.
  • Define task objectives and reward functions.

Skills

Python programming
Reinforcement Learning
Data evaluation
Data scraping

Job description

We are looking for an RL Environment Data Engineer / Researcher to design, build, and refine reinforcement learning training environments across different domains. This role will focus on data collection, task definition, reward design, evaluation criteria, anti-reward-hacking mechanisms, and post-training validation of environment data effectiveness.

Responsibilities
  • - Design and improve RL training environments across various task domains.
  • - Collect, clean, structure, and evaluate data used for RL environment construction and model post-training.
  • - Define task objectives, reward functions, and evaluation standards to ensure reliable and reproducible training signals.
  • - Develop technical approaches to prevent reward hacking and identify loopholes in reward design.
  • - Build validation environments to assess the effectiveness of post-training data and RL environment design.
  • - Collaborate with research, engineering, and data teams to improve environment coverage, task difficulty, and evaluation reliability.
  • - Follow research progress in RL environments, data evaluation, AI agents, and post-training methods, and apply relevant findings to production workflows.
Requirements
  • - Strong coding skills, especially in Python, with the ability to independently build data pipelines, environments, and evaluation tools.
  • - Proficiency with AI coding tools for code generation, debugging, refactoring, and rapid experimentation.
  • - Solid understanding of reinforcement learning, post-training, reward function design, environment design, and data evaluation.
  • - Ability to translate real-world tasks into trainable and measurable RL environments.
  • Experience with data scraping, data cleaning, annotation, or data quality assessment is preferred.
  • - Experience with LLM agents, RLHF/RLAIF, coding agents, automated evaluation, or benchmark construction is a strong plus.
  • - Strong experimental mindset and engineering execution, with the ability to continuously improve systems based on data and evaluation results.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Environment Architect & Data Research Engineer
RL Environment Architect & Data Research Engineer

Eigent AI • Greater London

Hybrid
GBP 65,000 - 85,000
Staff AI Engineer: RL Environments & Agent Training
Staff AI Engineer: RL Environments & Agent Training

United States Digital Space LLC • Greater London

On-site
GBP 70,000 - 110,000
Weekly lunch stipend
Full health and dental benefits
RRSP matching/401K
Remote RL Environments Engineer
Remote RL Environments Engineer

Cohere • Greater London

Hybrid
GBP 70,000 - 120,000
Weekly lunch stipend
Full health and dental benefits
RRSP matching / Pension
+6
Software Engineer, RL Data
Software Engineer, RL Data

United States Digital Space LLC • Greater London

Hybrid
GBP 242,000 - 368,000
Generous vacation and parental leave
Flexible working hours
Office space for collaboration
Member of Technical Staff - Research Scientist
Member of Technical Staff - Research Scientist

General Reasoning, Inc. • Greater London

On-site
GBP 70,000 - 90,000
Research Engineer, RL Scaling Science
Research Engineer, RL Scaling Science

Anthropic • Greater London

Hybrid
GBP 375,000 - 640,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
+1
Machine Learning Engineer- Reinforcement Learning
Machine Learning Engineer- Reinforcement Learning

Wave Recruitment • Greater London

Hybrid
GBP 75,000 - 120,000
Hybrid work
Visa sponsorship available
Direct access to CTO
Senior Research Scientist, RL & Post-Training (Hybrid)
Senior Research Scientist, RL & Post-Training (Hybrid)

Deepl-Se • Greater London

Hybrid
GBP 90,000 - 130,000
Hybrid work
Hack Fridays
30 days annual leave
+2
Remote AI Agent Environments Engineer
Remote AI Agent Environments Engineer

AI Chopping Block • Greater London

Hybrid
GBP 110,000 - 170,000
Lunch stipend
Health and dental benefits
RRSP matching
+5
Research Engineer, RL Scaling Science New London, UK
Research Engineer, RL Scaling Science New London, UK

Alcides Fonseca • Greater London

Hybrid
GBP 70,000 - 90,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours