Reinforcement Learning Lead

Datamentors

Portugal

Presencial

EUR 90 000 - 130 000

Tempo integral

Há 2 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Recebe uma resposta deste empregador — um currículo e uma carta de apresentação adaptados exatamente ao que estão a contratar.

Ultrapassa os filtros ATS

Resumo da oferta

Datamentors is building the next generation of autonomous robotics, combining AI, robotics, and VLA models. You will own the post-training stage of the Ardia humanoid platform, raising success rate and robustness through RL fine-tuning, reward modelling, and sim-to-real loops.

In this hands-on leadership role you will set the RL roadmap, design training and evaluation infrastructure, and mentor a small team to turn demonstrations into dependable production-ready behavior.

Qualificações

  • Strong RL expertise with policy-gradient methods and offline RL.
  • Hands-on robot learning/embodied AI experience with VLA models.
  • Experience with sim-to-real and at least one major simulator (Isaac Gym/MuJoCo).
  • Solid ML engineering with PyTorch and distributed training.
  • Evidence of shipping policies from demo to production.

Responsabilidades

  • Post‑VLA RL fine‑tuning — design and run the pipeline that takes a pretrained/imitation‑learned VLA policy and improves it with reinforcement learning (online and offline RL, RLHF/RL‑from‑feedback, preference and reward‑model approaches) to lift task success rates and reduce failure modes.
  • Reward design & evaluation — define reward signals, success criteria, and automated evaluation harnesses that actually correlate with real‑world manipulation and whole‑body performance, and that catch regressions before deployment.
  • Sim‑to‑real — build and tune the simulation‑to‑hardware transfer loop (domain randomization, residual policies, real‑world fine‑tuning) so gains in simulation hold up on the physical robot.
  • Data flywheel — establish the loop that turns robot rollouts and teleop data into better policies: autonomous data collection, filtering, labelling, and continual retraining.
  • Integration across the stack — work with the teams owning the orchestration layer, the VLA policy, and whole‑body control so RL improvements compose cleanly rather than fighting the rest of the system.
  • Technical leadership — set the RL roadmap, make build‑vs‑adopt calls on frameworks and tooling, mentor engineers, and represent the work to the wider team and external partners.

Conhecimentos

RL expertise
Robotics AI
Sim-to-real
ML engineering
Policy optimization
Shipping products

Ferramentas

PyTorch
MuJoCo
Isaac Gym

Descrição da oferta de emprego

About Datamentors Datamentors is building the next generation of autonomous robotics through a combination of advanced AI, robotics, and proprietary Vision‑Language‑Action (VLA) models. Our mission is to create intelligent robotic systems capable of understanding natural language, reasoning about the world, and acting autonomously in complex environments.

We are assembling a world‑class engineering team to tackle some of the hardest challenges in robotics, AI, perception, and autonomy. The engineers who join us today will have a direct impact on the core technologies that power Ardia and the future of intelligent machines

Role Overview Own the post‑training stage of the Ardia humanoid platform — taking pretrained, behaviour‑cloned Vision‑Language‑Action policies and systematically lifting their success rate, precision, and robustness through RL fine‑tuning, reward modelling, and a tight sim‑to‑real loop. This hands‑on leadership role sets the technical direction for RL across the platform, builds the training and evaluation infrastructure, and grows a small team around it. You'll be the person who turns a capable‑but‑inconsistent humanoid into a dependable one. The difference between a research demo and a product.

Responsabilities
  • Post‑VLA RL fine‑tuning — design and run the pipeline that takes a pretrained/imitation‑learned VLA policy and improves it with reinforcement learning (online and offline RL, RLHF/RL‑from‑feedback, preference and reward‑model approaches) to lift task success rates and reduce failure modes.
  • Reward design & evaluation — define reward signals, success criteria, and automated evaluation harnesses that actually correlate with real‑world manipulation and whole‑body performance, and that catch regressions before deployment.
  • Sim‑to‑real — build and tune the simulation‑to‑hardware transfer loop (domain randomization, residual policies, real‑world fine‑tuning) so gains in simulation hold up on the physical robot.
  • Data flywheel — establish the loop that turns robot rollouts and teleop data into better policies: autonomous data collection, filtering, labelling, and continual retraining.
  • Integration across the stack — work with the teams owning the orchestration layer, the VLA policy, and whole‑body control so RL improvements compose cleanly rather than fighting the rest of the system.
  • Technical leadership — set the RL roadmap, make build‑vs‑adopt calls on frameworks and tooling, mentor engineers, and represent the work to the wider team and external partners.
Requirements
  • Strong, demonstrable RL expertise: PPO/GRPO and related policy‑gradient methods, offline RL, RLHF/preference‑based RL, and reward modelling — you understand where each breaks and why.
  • Robot learning / embodied AI experience — ideally hands‑on with VLA models (e.G. OpenVLA, π0, GR00T‑class, or comparable manipulation/locomotion policies).
  • Practical sim‑to‑real experience and fluency with at least one major simulator (Isaac Lab / Isaac Gym, MuJoCo, or similar).
  • Solid ML engineering: PyTorch, distributed training, GPU‑efficient pipelines, and the discipline to build reproducible experiments and rigorous evaluation.
  • Evidence of shipping — you've taken a policy from "works in a demo" to "works reliably," not just published benchmarks.
Nice to have
  • Direct experience with humanoid or whole‑body control (locomotion + manipulation), and with the realities of training on real hardware.
  • Familiarity with VLA post‑training specifically — fine‑tuning, distillation, or RL on top of large pretrained action models.
  • Comfort working close to perception (vision, point clouds) and to low‑level control.
  • Open‑source contributions to robot learning or RL frameworks.
  • Experience standing up data‑collection / teleoperation pipelines.
  • Publications or applied work at the VLA / robot‑learning frontier (ICRA / CoRL / RSS‑level).
Why Datamentors

Own the post‑training layer end to end, with the autonomy to build it your way.

The architecture, hardware, and base policies are already in place — your RL work is the leap that turns a capable demo into a dependable product.

Hands‑on technical leadership: set the RL roadmap, make the tooling calls, and grow a small team around the work.

We assess candidates on demonstrated ability, not credentials alone — if you've done the work and can show it, we want to talk.

From research demo to dependable product — own it.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Reinforcement Learning Lead
Reinforcement Learning Lead

Datamentors • Região Autónoma Da Madeira

Presencial
EUR 120 000 - 180 000
Reinforcement Learning Lead
Reinforcement Learning Lead

Datamentors • Lisboa

Presencial
EUR 90 000 - 130 000
RL Lead, Humanoid Robotics & Post-Training
RL Lead, Humanoid Robotics & Post-Training

Datamentors • Portugal

Presencial
EUR 90 000 - 130 000
Head of RL for Humanoid Autonomy
Head of RL for Humanoid Autonomy

Datamentors • Lisboa

Presencial
EUR 90 000 - 130 000
Lead RL Engineer for Humanoid Robotics
Lead RL Engineer for Humanoid Robotics

Datamentors • Região Autónoma Da Madeira

Presencial
EUR 120 000 - 180 000
Software Engineer — Robotics/AI
Software Engineer — Robotics/AI

Career Genie, Inc • Lisboa

Híbrido
EUR 40 000 - 60 000
Competitive salary with relocation support
Equity participation
Direct access to real robots
Senior AI Engineer — Spatial Intelligence, Segmentation & 3D
Senior AI Engineer — Spatial Intelligence, Segmentation & 3D

Career Genie, Inc • Lisboa

Presencial
EUR 45 000 - 70 000
Relocation support
Housing assistance
Teleoperator
Teleoperator

Career Genie, Inc • Portugal

Híbrido
EUR 30 000 - 50 000
Equity participation
Relocation support
Competitive compensation
+1
Senior Electronics Engineer
Senior Electronics Engineer

Datamentors • Região Autónoma Da Madeira

Presencial
EUR 60 000 - 90 000
Health insurance
Relocation/housing support for hires
IT / System Admin & Ops
IT / System Admin & Ops

Career Genie, Inc • Portugal

Teletrabalho
EUR 30 000 - 50 000
Equity participation
Relocation support
Competitive compensation
+1