Head of RL for Humanoid Autonomy

Datamentors

Lisboa

Presencial

EUR 90 000 - 130 000

Tempo integral

Há 2 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Destaca-te para esta função — gera um currículo e uma carta de apresentação personalizados em cerca de um minuto.

Ultrapassa os filtros ATS

Resumo da oferta

Datamentors seeks an Reinforcement Learning Lead to own the post-training stage of the Ardia humanoid platform, lifting success, precision and robustness through RL fine-tuning, reward modelling and a tight sim-to-real loop. You will set the technical direction, build training/evaluation infra, and grow a small team around it.

You will lead the RL roadmap, mentor engineers, and collaborate with perception and control teams to turn demonstrations into a dependable product.

Qualificações

  • Strong, demonstrable RL expertise: PPO/GRPO and related policy-gradient methods, offline RL, RLHF/preference-based RL, and reward modelling.
  • Robot learning / embodied AI experience - ideally hands-on with VLA models.
  • Practical sim-to-real experience and fluency with major simulators (Isaac Lab / Isaac Gym, MuJoCo).
  • Solid ML engineering: PyTorch, distributed training, GPU-efficient pipelines, reproducible experiments.

Responsabilidades

  • Post-VLA RL fine-tuning - design and run the pipeline that takes a pretrained/imitation-learned VLA policy and improves it with reinforcement learning (online and offline RL, RLHF/RL-from-feedback, preference and reward-model approaches) to lift task success rates and reduce failure modes.
  • Reward design & evaluation - define reward signals, success criteria, and automated evaluation harnesses that actually correlate with real-world manipulation and whole-body performance, and that catch regressions before deployment.
  • Sim-to-real - build and tune the simulation-to-hardware transfer loop (domain randomization, residual policies, real-world fine-tuning) so gains in simulation hold up on the physical robot.
  • Data flywheel - establish the loop that turns robot rollouts and teleop data into better policies: autonomous data collection, filtering, labelling, and continual retraining.
  • Integration across the stack - work with the teams owning the orchestration layer, the VLA policy, and whole-body control so RL improvements compose cleanly rather than fighting the rest of the system.
  • Technical leadership - set the RL roadmap, make build-vs-adopt calls on frameworks and tooling, mentor engineers, and represent the work to the wider team and external partners.

Conhecimentos

RL expertise
Embodied AI
Sim-to-real
PPO/GRPO
Reward modelling

Ferramentas

PyTorch
Isaac Gym
MuJoCo
GPUs & distributed

Descrição da oferta de emprego

Datamentors seeks an Reinforcement Learning Lead to own the post-training stage of the Ardia humanoid platform, lifting success, precision and robustness through RL fine-tuning, reward modelling and a tight sim-to-real loop. You will set the technical direction, build training/evaluation infra, and grow a small team around it.

You will lead the RL roadmap, mentor engineers, and collaborate with perception and control teams to turn demonstrations into a dependable product.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Lead RL Engineer for Humanoid Robotics
Lead RL Engineer for Humanoid Robotics

Datamentors • Região Autónoma Da Madeira

Presencial
EUR 120 000 - 180 000
RL Lead, Humanoid Robotics & Post-Training
RL Lead, Humanoid Robotics & Post-Training

Datamentors • Portugal

Presencial
EUR 90 000 - 130 000
Reinforcement Learning Lead
Reinforcement Learning Lead

Datamentors • Lisboa

Presencial
EUR 90 000 - 130 000
Reinforcement Learning Lead
Reinforcement Learning Lead

Datamentors • Região Autónoma Da Madeira

Presencial
EUR 120 000 - 180 000
Reinforcement Learning Lead
Reinforcement Learning Lead

Datamentors • Portugal

Presencial
EUR 90 000 - 130 000
Robotics RL Lead: From Lab to Production
Robotics RL Lead: From Lab to Production

Career Genie, Inc • Portugal

Teletrabalho
EUR 70 000 - 90 000
Senior Robotics Electronics Engineer - Hardware Lead
Senior Robotics Electronics Engineer - Hardware Lead

Datamentors • Região Autónoma Da Madeira

Presencial
EUR 60 000 - 90 000
Health insurance
Relocation/housing support for hires
Software Engineer — Robotics/AI
Software Engineer — Robotics/AI

Career Genie, Inc • Lisboa

Híbrido
EUR 40 000 - 60 000
Competitive salary with relocation support
Equity participation
Direct access to real robots
Senior AI Engineer — Spatial Intelligence, Segmentation & 3D
Senior AI Engineer — Spatial Intelligence, Segmentation & 3D

Career Genie, Inc • Lisboa

Presencial
EUR 45 000 - 70 000
Relocation support
Housing assistance
Senior Electronics Engineer
Senior Electronics Engineer

Datamentors • Região Autónoma Da Madeira

Presencial
EUR 60 000 - 90 000
Health insurance
Relocation/housing support for hires