PhD Studentship: Causal Reinforcement Learning

Phaidra

Cambridge

On-site

GBP 21,000 - 26,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

MacBook
Unlimited PTO
Parental leave
Equity
Competitive compensation

Job summary

Phaidra invites applications for a fully funded 4‑year PhD studentship hosted at the University of Cambridge. You will explore integrating causal reasoning into reinforcement learning to build robust AI control systems for industrial settings.

The project covers theoretical foundations, algorithm development, and benchmarking, with supervision from Cambridge and Phaidra. International applicants welcome; January 2027 start.

Qualifications

  • A first-class or upper second-class honours degree (or equivalent) in Computer Science, Mathematics, Engineering, Statistics, or a related technical field.
  • Strong background in at least one of: reinforcement learning, machine learning, probabilistic modelling, or control theory.
  • Proficiency in Python and standard ML libraries (PyTorch, NumPy, SciPy, scikit‑learn).
  • Clear scientific writing skills and the ability to communicate research to both academic and applied audiences.

Responsibilities

  • Contribute to theoretical foundations of offline and online RL with a causal lens.
  • Develop algorithms that leverage causal structure to improve generalisation and policy guarantees.
  • Benchmark proposed methods in controlled simulated environments with known causal structure.

Skills

Python programming
Reinforcement learning
Machine learning
Probabilistic modelling
Scientific writing

Education

First-class or upper second-class honours degree in CS/Math/Engineering/Stats

Tools

PyTorch
NumPy
SciPy
scikit-learn

Job description

About Phaidra

Phaidra is building the future of industrial automation.

The world today is filled with static, monolithic infrastructure. Factories, power plants, buildings, etc. operate the same they've operated for decades — because the controls programming is hard-coded. Thousands of lines of rules and heuristics that define how the machines interact with each other. The result of all this hard‑coding is that facilities are frozen in time, unable to adapt to their environment while their performance slowly degrades.

Phaidra creates AI‑powered control systems for the industrial sector, enabling industrial facilities to automatically learn and improve over time. Specifically:

  • We use reinforcement learning algorithms to provide this intelligence, converting raw sensor data into high‑value actions and decisions.
  • We focus on industrial applications, which tend to be well‑sensorised with measurable KPIs — perfect for reinforcement learning.
  • We enable domain experts (our users) to configure the AI control systems (i.e. agents) without writing code. They define what they want their AI agents to do, and we do it for them.

Our team has a track record of applying AI to some of the toughest problems. From achieving superhuman performance with DeepMind's AlphaGo, to reducing the energy required to cool Google’s Data Centers by 40%, we deeply understand AI and how to apply it in production for massive impact.

Phaidra’s ability to achieve its mission is determined by our ability to work together — as defined by our core values: Transparency, Collaboration, Operational Excellence, Ownership, and Empathy. We seek individuals who embody these values, as they are instrumental in ensuring our team consistently delivers excellence and fosters an engaging and supportive culture.

Phaidra is based in the USA, but we are 100% remote with no physical office. We hire employees internationally with the help of our partner OysterHR. Our team is currently located throughout the USA, Canada, UK, Sweden, Spain, Portugal, the Netherlands, Singapore, Australia, and India.

About the Project

Phaidra builds autonomous AI control systems for data centre and industrial infrastructure. We deploy reinforcement learning in production on some of the world's most complex physical systems. The hard, unsolved research problems are the same ones that matter in practice. This PhD project is an opportunity to work on foundational RL research while staying grounded in real‑world challenges.

Reinforcement Learning (RL) has emerged as a powerful framework for sequential decision‑making. Yet a fundamental limitation remains: agents trained on historical data under fixed policies often exploit spurious correlations that break at deployment time, especially when the environment shifts or the new policy explores previously unseen regions of the state‑action space.

This PhD project tackles that limitation by integrating causal reasoning into RL. Causal inference provides a formal language (causal graphs, interventional queries, counterfactuals) for distinguishing stable structural relationships from incidental correlations. The research will investigate how these tools can make RL agents more robust and generalisable, particularly in real‑world industrial settings.

The project will proceed in three phases:

  1. Theoretical Foundations: formalising policy learning from biased, small datasets through a causal lens; characterising how confounding and mediators affect offline RL.
  2. Algorithm Development: building RL algorithms that leverage known or learned causal structure to improve out‑of‑distribution generalisation and provide policy guarantees.
  3. Benchmarking & Evaluation: evaluating proposed methods on controlled simulated environments with known causal structure, benchmarked against standard and offline RL baselines.
Supervisors
  • Academic Supervisor: Prof. Alessandro Abate, Department of Engineering, University of Cambridge
  • Industrial Co‑supervisors: Dr. Miguel Suau and Dr. Alec Edwards, Phaidra

The student will be based primarily at the University of Cambridge, with the opportunity to spend time at Phaidra.

Funding & Duration

This is a fully funded 4‑year PhD studentship, expected to start January 2027, co‑funded by Phaidra and administered by the University of Cambridge.

Who You Are

You are a curious and technically rigorous researcher who wants to work at the intersection of causal inference and sequential decision‑making. You are excited by foundational questions with real‑world stakes and want your PhD to contribute both to the academic literature and to the practical deployment of intelligent systems.

Key Qualifications
  • A first‑class or upper second‑class honours degree (or equivalent) in Computer Science, Mathematics, Engineering, Statistics, or a related technical field.
  • Strong background in at least one of: reinforcement learning, machine learning, probabilistic modelling, or control theory.
  • Proficiency in Python and standard ML libraries (PyTorch, NumPy, SciPy, scikit‑learn).
  • Clear scientific writing skills and the ability to communicate research to both academic and applied audiences.
  • Eligibility to study at the University of Cambridge (international students welcome; English language requirements apply).
Preferred Skills & Experience
  • Familiarity with causal inference, causal graphical models, or structural equation models.
  • Prior research experience (undergraduate thesis, MSc dissertation, research internship, or publications).
  • Experience with offline RL, batch RL, or safe RL.
  • Exposure to applying ML to real‑world physical or industrial systems.
Benefits & Perks
  • Fast‑paced, team‑oriented environment where your work directly shapes the company’s direction.
  • We are a 100% remote company.
  • Competitive compensation & meaningful equity.
  • Outsized responsibilities & professional development.
  • Training is foundational; functional, customer immersion, and development training.
  • Medical, dental, and vision insurance (exact benefits vary by region).
  • Unlimited paid time off, with a required minimum of 20 days per year.
  • Paid parental leave (exact benefits vary by region).
  • Flexible stipends to support your workspace, well‑being, and continued professional development.
  • Company MacBook.
Equal Opportunity Employment

Phaidra is an Equal Opportunity Employer; employment with Phaidra is governed on the basis of merit, competence, and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability, or any other legally protected status. We welcome diversity and strive to maintain an inclusive environment for all employees. If you need assistance with completing the application process, please contact us at hiring@phaidra.ai.

E‑Verify Notice

Phaidra participates in E‑Verify, an employment authorization database provided through the U.S. Department of Homeland Security (DHS) and Social Security Administration (SSA). As required by law, we will provide the SSA and, if necessary, the DHS, with information from each new employee’s Form I‑9 to confirm work authorization for those residing in the United States.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

PhD in Causal RL for Industrial Automation (Remote)
PhD in Causal RL for Industrial Automation (Remote)

Phaidra • Cambridge

On-site
GBP 21,000 - 26,000
MacBook
Unlimited PTO
Parental leave
+2
Research Engineer, Machine Learning (Reinforcement Learning) London, UK
Research Engineer, Machine Learning (Reinforcement Learning) London, UK

Alcides Fonseca • Greater London

Hybrid
GBP 89,000 - 133,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Software Engineer, RL Data
Software Engineer, RL Data

United States Digital Space LLC • Greater London

Hybrid
GBP 242,000 - 368,000
Generous vacation and parental leave
Flexible working hours
Office space for collaboration
Research Engineer, RL Scaling Science New London, UK
Research Engineer, RL Scaling Science New London, UK

Alcides Fonseca • Greater London

Hybrid
GBP 70,000 - 90,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Senior Model-Based RL Scientist (Remote UK)
Senior Model-Based RL Scientist (Remote UK)

Phaidra • United Kingdom

Remote
GBP 85,000 - 117,000
Equity
Remote-first company
Medical, dental, and vision insurance
+4
Senior AI Engineer
Senior AI Engineer

Causaly • Greater London

On-site
GBP 90,000 - 130,000
Competitive compensation package
Private medical & dental insurance
Life insurance (4 x salary)
+5
PhD Graduate AI & Algorithms Engineer (2026 start)
PhD Graduate AI & Algorithms Engineer (2026 start)

Cambridge Consultants • Cambridge

On-site
GBP 45,000 - 65,000
AI Research Engineer - Reinforcement Learning
AI Research Engineer - Reinforcement Learning

Helsing • Greater London

On-site
GBP 60,000 - 78,000
Competitive salary
Stock options (ESOP)
Relocation support
+6
Technical Specialist - Multi-Agent Security
Technical Specialist - Multi-Agent Security

Advanced Research & Invention Agency • Greater London

On-site
GBP 70,000 - 105,000
27 days annual leave
Hybrid working arrangements
Learning and development opportunities
+5
Member of Technical Staff (Post Training)
Member of Technical Staff (Post Training)

Inherentlabs • Greater London

On-site
GBP 80,000 - 100,000