Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Humanoid is building humanoid robots and software to operate in real-world environments. This internship offers hands-on experience across our research stack in London, with mentorship from experienced researchers and engineers.
You will work on reinforcement learning, world models, and inference & optimisation, with opportunities to contribute to real systems from early on. We welcome master's or PhD students in CS/ML/robotics who are proactive and eager to learn in a fast-moving,
Here at Humanoid, we believe in a future where robots amplify human potential. That’s why we’ve set out on a mission to build the world’s most capable, commercially-scalable, and safe humanoid robots. We’re bringing that mission to life with HMND‑01 - our rapidly developed humanoid platform being deployed in real industrial environments - and we’re growing the team to take it even further.
We're building software systems that enable robots to operate effectively in the real world, expanding human capability and redefining how work gets done.
We're looking for interns who are curious, proactive, and excited to work on real-world robotic systems.
Depending on your interests and skills, you will be able to work across our research stack: reinforcement learning, world models, pretraining, and inference & optimisation. That spans everything from training policies in simulation, through building the generative models that let robots predict their world, to squeezing models onto real-time edge compute. You'll collaborate closely with the team to find where you can have the most impact, and we're looking for people who are excited to dive into unfamiliar areas and learn quickly.
This is a full-time internship (5 days per week), based in our London office, where you'll contribute to real systems from early on with guidance and support from experienced researchers and engineers.
Duration: 12 to 24 weeks | Start date: Flexible | Compensation: Competitive pay and perks
Reinforcement Learning
Train language-vision conditioned manipulation policies via RL in the real world
Construct challenging and diverse suites of manipulation tasks and RL models in simulation (Isaac Sim, MuJoCo)
Experiment with ways of bringing policies trained in simulation to the real world
World Models
Action-conditioned video prediction and dynamics models that stay physically consistent over long horizons
Use world models as learned simulators: score candidate policies offline and generate synthetic rollouts for training
Build fidelity metrics that quantify where the world model can be trusted
Pre- and post-training
In-context learning
Short and long term memory
Post-training VLA models on specific production-grade use cases
Different data modalities, closing embodiment gap between human and robot data, data diversity and attribution.
Inference & Optimisation
Optimise models for real-time edge inference on robot hardware: profiling, quantisation, and latency/throughput trade-offs
Improve training and data-loading performance across distributed GPU infrastructure
Candidates pursuing or holding a master’s or PhD in computer science, machine learning, robotics, or a related field.
Strong foundations in machine learning; strong Python and hands-on experience with PyTorch or JAX.
Interest in one or more of: reinforcement learning, world models and generative video, VLA/multimodal models, or ML systems and inference optimisation.
Experience running experiments and interpreting results with rigour.
Ability to take ownership and iterate with guidance.
Strong problem-solving skills and attention to detail.
Fast learner, comfortable in a research-driven, fast-moving environment.
Free daily breakfast, catered lunch, and snacks in-office.
Work at the frontier - collaborate daily with world-class engineers, researchers, and product experts building the next generation of AI and humanoid robotics.
Real ownership - direct access to founding leadership, meaningful input on product direction, and the ability to drive key initiatives from day one.