Machine Learning Engineer- Reinforcement Learning

Wave Recruitment

Greater London

Hybrid

GBP 75,000 - 120,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work
Visa sponsorship available
Direct access to CTO

Job summary

Wave Recruitment is helping a client in London recruit an ML Engineer specializing in reinforcement learning to design and deploy agents for live data-centre cooling. You will work across research and deployment, reporting to the CTO/Head of AI, and balance experimentation with producing stable production code on hybrid sites.

The role requires 3–5 years of RL experience, Python, and experience with PyTorch or JAX, plus a strong physics or engineering background.

Qualifications

  • 3–5 years of experience training and deploying deep RL agents in Python.
  • Experience with PyTorch or JAX and RL libraries such as Gymnasium.
  • Background in physical systems (mechanical, electrical, controls) and ability to reason about physical feasibility.

Responsibilities

  • Train and deploy deep RL agents for live cooling control.
  • Design reward functions and constraints aligned with ASHRAE standards and SLAs.
  • Bridge research exploration and engineering to run on live sites.
  • Develop and maintain simulators and digital twins.

Skills

Python
Deep RL
Research-to-Deployment
System thinking

Education

Engineering/Physics background

Tools

PyTorch
JAX
Gymnasium

Job description

ML Engineer - Reinforcement Learning London (hybrid, 1 day/week in Kings Cross)- Solve Data Centres Cooling issues

Cooling is one of the largest items on a data centre's energy bill, and most sites run it conservatively because getting it wrong puts the hardware at risk. Our client trains reinforcement learning agents to control cooling systems on live sites, cutting cooling energy without breaching the temperature and humidity limits operators are contractually bound to.

They're hiring an ML Engineer - Reinforcement Learning to build those agents and get them running on real data centres. You'll report to the CTO / Head of AI and work across the line between research and deployment.

The System

The agents don't learn on the live plant. They train against a digital twin of each site, then move to production once they're safe.

  • Reward and constraint design is shaped by ASHRAE standards and customer SLAs - air temperature, humidity, and rate-of-change limits on cooling air and chilled water setpoints
  • Training is federated across multiple sites. Agents share learned control strategies without any site's operational data leaving the building, which delivers significantly more savings than a single-site approach
  • Models are deployed on-prem at the edge, then monitored and retrained in place
What You'll Own
  • Train and deploy deep RL agents for live cooling control
  • Design reward functions and constraints that hold up against physical limits and SLAs, not just in a notebook
  • Move between research-style exploration and the engineering work to make something stable on a real site

Simulation and Digital Twins

  • Build and improve the physics-based simulators, surrogate models, and digital twins the agents train against
  • Close the gap between what works in simulation and what holds on real hardware

Production and Deployment

  • Federated and distributed training across sites
  • Edge deployment, monitoring, and retraining of agents already running in production
What We're Looking For
  • 3-5 years training and deploying deep RL agents in Python
  • PyTorch or JAX, and RL libraries such as Gymnasium
  • A background in physical systems - engineering (mechanical, electrical, structural, biomedical), physics, robotics, autonomous driving, or control systems - and the instinct to reason about what's physically possible, not only what's mathematically possible
  • Comfortable iterating between research exploration and the engineering needed to run on a live site

Useful

  • Control systems (classical control, MPC), or HVAC, thermodynamics, power systems, or data centre operations
  • Federated learning, distributed training, or edge ML deployment
  • Simulation experience - building or using physics-based simulators, digital twins, surrogate models, or large physics models
  • Published research or open-source contributions
Who You Are

You want both halves of this job. You'll run experiments and read papers, but you also want your work controlling real equipment, with the constraints that come with that. RL experience limited to advertising or multi-armed bandits won't carry over here - the physical world doesn't behave like a recommendation system. A pure maths or CS background with no feel for physical systems will struggle, and so will anyone after a pure research seat or a pure production one.

This sits in the middle.

What's On Offer
  • A genuine technical problem: RL on physical systems, under real constraints, deployed on live infrastructure
  • Direct access to the CTO and founding team
  • Hybrid working, one day a week in the Kings Cross office
  • Visa sponsorship available on a case-by-case basis
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Edge RL Engineer for Live Data Centre Cooling
Edge RL Engineer for Live Data Centre Cooling

Wave Recruitment • Greater London

Hybrid
GBP 75,000 - 120,000
Hybrid work
Visa sponsorship available
Direct access to CTO
Applied ML Scientist
Applied ML Scientist

Oliver Bernard • Greater London

Hybrid
GBP 80,000 - 110,000
Reinforcement Learning Researcher
Reinforcement Learning Researcher

AI Startups UK • Greater London

Hybrid
GBP 90,000 - 140,000
Machine Learning Engineer (Autonomous Systems)
Machine Learning Engineer (Autonomous Systems)

Understanding Recruitment • City Of London

Hybrid
GBP 70,000 - 110,000
Research Engineer, RL Scaling Science
Research Engineer, RL Scaling Science

Anthropic • Greater London

Hybrid
GBP 375,000 - 640,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
+1
Simulation Engineer - Manipulation
Simulation Engineer - Manipulation

Thehumanoid • Greater London

On-site
GBP 60,000 - 80,000
23 days annual leave
Fully funded private healthcare
Equity options
+3
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Machine Learning Engineer (AI Start-Up) - Multiple Roles & Differing Seniorities - £70k - £110k
Machine Learning Engineer (AI Start-Up) - Multiple Roles & Differing Seniorities - £70k - £110k

Few&Far • Greater London

On-site
GBP 70,000 - 110,000
Private medical insurance
Equity in business growth
Downtown London office culture
+1
Reinforcement Learning (RL) Engineer, Manipulation - Full UK Visa Sponsorship Available
Reinforcement Learning (RL) Engineer, Manipulation - Full UK Visa Sponsorship Available

EasyInfoBlog.com LLC • Greater London

On-site
GBP 80,000 - 120,000
Family‑First Relocation
Full UK Visa Sponsorship
Fully Catered Meals
+1
Research Engineer, Pretraining Scaling - London
Research Engineer, Pretraining Scaling - London

Anthropic • Greater London

On-site
GBP 250,000 - 435,000
Equity benefits
Visa sponsorship