Reinforcement Learning Lead

Datamentors

Região Autónoma Da Madeira

Presencial

EUR 120 000 - 180 000

Tempo integral

14 dias+
Gerador de candidaturas

Uma candidatura feita para esta oferta — um currículo e uma carta de apresentação personalizados que vão ao encontro do anúncio.

Ultrapassa os filtros ATS

Resumo da oferta

Datamentors is seeking a leadership-level engineer to own the post-training RL layer for the Ardia humanoid platform. You will lift success rates, robustness, and precision through RL fine-tuning, reward modelling, and a tight sim-to-real loop, shaping the technical direction for RL across the platform.

You will oversee training infrastructure, evaluation, and a growing team, turning capable demonstrations into dependable product capabilities with close collaboration across teams.

Qualificações

  • Strong RL expertise including PPO/GRPO, offline RL, RLHF, and reward modeling.
  • Hands-on experience with VLA models or comparable manipulation/locomotion policies.
  • Practical sim-to-real experience with major simulators.

Responsabilidades

  • Design and run RL post-training pipelines to improve VLA policies.
  • Define reward signals and automated evaluation harnesses for real-world performance.
  • Build the sim-to-real transfer loop and tune domain randomization.
  • Establish data flywheel for autonomous data collection and retraining.
  • Coordinate integration with orchestration, VLA policy, and entire stack.
  • Provide technical leadership and mentor engineers, roadmap decisions.

Conhecimentos

RL expertise
Embodied AI
Sim-to-Real experience
ML engineering
Shipping demonstrated results

Formação académica

PhD in ML/Robotics

Ferramentas

PyTorch
MuJoCo
Isaac Gym

Descrição da oferta de emprego

About Datamentors

Datamentors is building the next generation of autonomous robotics through a combination of advanced AI, robotics, and proprietary Vision-Language-Action (VLA) models. Our mission is to create intelligent robotic systems capable of understanding natural language, reasoning about the world, and acting autonomously in complex environments.

We are assembling a world‑class engineering team to tackle some of the hardest challenges in robotics, AI, perception, and autonomy. The engineers who join us today will have a direct impact on the core technologies that power Ardia and the future of intelligent machines

Role Overview

Own the post‑training stage of the Ardia humanoid platform — taking pretrained, behaviour‑cloned Vision‑Language‑Action policies and systematically lifting their success rate, precision, and robustness through RL fine‑tuning, reward modelling, and a tight sim‑to‑real loop. This hands‑on leadership role sets the technical direction for RL across the platform, builds the training and evaluation infrastructure, and grows a small team around it. You’ll be the person who turns a capable‑but‑inconsistent humanoid into a dependable one. The difference between a research demo and a product.

Responsabilities
  • Post‑VLA RL fine‑tuning — design and run the pipeline that takes a pretrained/imitation‑learned VLA policy and improves it with reinforcement learning (online and offline RL, RLHF/RL‑from‑feedback, preference and reward‑model approaches) to lift task success rates and reduce failure modes.
  • Reward design & evaluation — define reward signals, success criteria, and automated evaluation harnesses that actually correlate with real‑world manipulation and whole‑body performance, and that catch regressions before deployment.
  • Sim‑to‑real — build and tune the simulation‑to‑hardware transfer loop (domain randomization, residual policies, real‑world fine‑tuning) so gains in simulation hold up on the physical robot.
  • Data flywheel — establish the loop that turns robot rollouts and teleop data into better policies: autonomous data collection, filtering, labelling, and continual retraining.
  • Integration across the stack — work with the teams owning the orchestration layer, the VLA policy, and whole‑body control so RL improvements compose cleanly rather than fighting the rest of the system.
  • Technical leadership — set the RL roadmap, make build‑vs‑adopt calls on frameworks and tooling, mentor engineers, and represent the work to the wider team and external partners.
Requirements
  • Strong, demonstrable RL expertise: PPO/GRPO and related policy‑gradient methods, offline RL, RLHF/preference‑based RL, and reward modelling — you understand where each breaks and why.
  • Robot learning / embodied AI experience — ideally hands‑on with VLA models (e.g. OpenVLA, π0, GR00T‑class, or comparable manipulation/locomotion policies).
  • Practical sim‑to‑real experience and fluency with at least one major simulator (Isaac Lab / Isaac Gym, MuJoCo, or similar).
  • Solid ML engineering: PyTorch, distributed training, GPU‑efficient pipelines, and the discipline to build reproducible experiments and rigorous evaluation.
  • Evidence of shipping — you’ve taken a policy from "works in a demo" to "works reliably," not just published benchmarks.
Nice to have
  • Direct experience with humanoid or whole‑body control (locomotion + manipulation), and with the realities of training on real hardware.
  • Familiarity with VLA post‑training specifically — fine‑tuning, distillation, or RL on top of large pretrained action models.
  • Comfort working close to perception (vision, point clouds) and to low‑level control.
  • Open‑source contributions to robot learning or RL frameworks.
  • Experience standing up data‑collection / teleoperation pipelines.
  • Publications or applied work at the VLA / robot‑learning frontier (ICRA / CoRL / RSS‑level).
Why Datamentors
  • Own the post‑training layer end to end, with the autonomy to build it your way.
  • The architecture, hardware, and base policies are already in place — your RL work is the leap that turns a capable demo into a dependable product.
  • Hands‑on technical leadership: set the RL roadmap, make the tooling calls, and grow a small team around the work.
  • We assess candidates on demonstrated ability, not credentials alone — if you’ve done the work and can show it, we want to talk.
From research demo to dependable product — own it.
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Reinforcement Learning Lead
Reinforcement Learning Lead

Datamentors • Portugal

Presencial
EUR 90 000 - 130 000
Reinforcement Learning Lead
Reinforcement Learning Lead

Datamentors • Lisboa

Presencial
EUR 90 000 - 130 000
RL Lead, Humanoid Robotics & Post-Training
RL Lead, Humanoid Robotics & Post-Training

Datamentors • Portugal

Presencial
EUR 90 000 - 130 000
Head of RL for Humanoid Autonomy
Head of RL for Humanoid Autonomy

Datamentors • Lisboa

Presencial
EUR 90 000 - 130 000
Lead RL Engineer for Humanoid Robotics
Lead RL Engineer for Humanoid Robotics

Datamentors • Região Autónoma Da Madeira

Presencial
EUR 120 000 - 180 000
Software Engineer — Robotics/AI
Software Engineer — Robotics/AI

Career Genie, Inc • Lisboa

Híbrido
EUR 40 000 - 60 000
Competitive salary with relocation support
Equity participation
Direct access to real robots
Teleoperator
Teleoperator

Career Genie, Inc • Portugal

Híbrido
EUR 30 000 - 50 000
Equity participation
Relocation support
Competitive compensation
+1
Senior AI Engineer — Spatial Intelligence, Segmentation & 3D
Senior AI Engineer — Spatial Intelligence, Segmentation & 3D

Career Genie, Inc • Lisboa

Presencial
EUR 45 000 - 70 000
Relocation support
Housing assistance
Senior Electronics Engineer
Senior Electronics Engineer

Datamentors • Região Autónoma Da Madeira

Presencial
EUR 60 000 - 90 000
Health insurance
Relocation/housing support for hires
IT / System Admin & Ops
IT / System Admin & Ops

Career Genie, Inc • Portugal

Teletrabalho
EUR 30 000 - 50 000
Equity participation
Relocation support
Competitive compensation
+1