Transformez ce poste en entretien — un CV et une lettre de motivation conçus selon ce que cet employeur recherche.
Built Different is hiring a founding junior RL engineer/researcher to help build RL environments and reward signals from day one. The role is hybrid across Amsterdam, London and Paris, with a competitive €140,000 base salary plus founding equity.
You have experience turning real human expertise into training signals and designing non-trivial RL reward structures, with a bias for fast iteration and first-principles thinking.
Founding Junior RL Engineer/Researcher (Mediors welcome also)
Hybrid | Amsterdam / London / Paris | €140,000 + Founding Equity
The best AI engineers aren't waiting for a job ad. They're waiting for the right moment. This might be it.
Anthropic leadership has discussed spending over $1 billion on RL environments in the next year alone. The market is moving - fast and at scale. One pre-seed company just closed a round at a valuation most Series A companies would envy, with frontier AI labs already as paying clients and acquisition interest already on the table (x2).
The reason? They cracked something the rest of the market hasn't.
Most teams build RL environments from synthetic data - easy to demo, easy to commoditise, brittle. This team mines real human behavioural data - how domain experts actually reason, decide, and solve complex tasks over long-horizon workflows. 10-100+ step environments. Closed-loop systems where environments, data, training, and evaluation are tightly integrated. Not proxies. Not short cuts.
25-30% uplift in model task success rates. 50-65% more training signals. Evals that reflect how humans actually work.
The funding just landed. Second time founders. The founding team is being built right now. There are very few seats. Are you going to be in one of them?
€140,000 + founding equity at a ~$40M seed valuation
Frontier AI labs as paying clients from day one
Environment design treated as a first-class problem - not an after thought
London / Paris hybrid with some remote options
You've built the evaluation-to-training loop - environments, rollouts, reward signals feeding post-training - at a serious RL or AI company
You fine-tuned foundation models, and done it on smaller open-weight models rather than just throwing scale at the problem
You think turning real human expertise into training signal is one of the most underrated problems in AI right now
Your RL is the genuinely hard kind - non-trivial reward design, long-horizon, multi-agent - not moving data from A to B
You move fast, think in first principles, and thrive without a playbook
You want to work at the forefront of innovation with frontier labs already paying to use what you build
Roles like this don't stay open.
Because great teams are Built Different.