Internship, Robot Learning Research
Quick Summary
score candidate poli
Here at Humanoid, we believe in a future where robots amplify human potential. That’s why we’ve set out on a mission to build the world’s most capable, commercially-scalable, and safe humanoid robots. We’re bringing that mission to life with HMND‑01 - our rapidly developed humanoid platform being deployed in real industrial environments - and we’re growing the team to take it even further.
We're building software systems that enable robots to operate effectively in the real world, expanding human capability and redefining how work gets done.
We're looking for interns who are curious, proactive, and excited to work on real-world robotic systems.
Depending on your interests and skills, you will be able to work across our research stack: reinforcement learning, world models, pretraining, and inference & optimisation. That spans everything from training policies in simulation, through building the generative models that let robots predict their world, to squeezing models onto real-time edge compute. You'll collaborate closely with the team to find where you can have the most impact, and we're looking for people who are excited to dive into unfamiliar areas and learn quickly.
What We Offer
~1 min readAction-conditioned video prediction and dynamics models that stay physically consistent over long horizons
Use world models as learned simulators: score candidate policies offline and generate synthetic rollouts for training
Build fidelity metrics that quantify where the world model can be trusted
In-context learning
Short and long term memory
Post-training VLA models on specific production-grade use cases
Different data modalities, closing embodiment gap between human and robot data, data diversity and attribution.
Optimise models for real-time edge inference on robot hardware: profiling, quantisation, and latency/throughput trade-offs
Improve training and data-loading performance across distributed GPU infrastructure
Candidates pursuing or holding a master’s or PhD in computer science, machine learning, robotics, or a related field.
Strong foundations in machine learning; strong Python and hands-on experience with PyTorch or JAX.
Interest in one or more of: reinforcement learning, world models and generative video, VLA/multimodal models, or ML systems and inference optimisation.
Experience running experiments and interpreting results with rigour.
Ability to take ownership and iterate with guidance.
Strong problem-solving skills and attention to detail.
Fast learner, comfortable in a research-driven, fast-moving environment.
Complete the challenge below and submit your solution as a public GitHub repository. You will be able to include your GitHub repository URL when you fill out the application form, alongside your name and CV. You have two weeks to complete the challenge and submit your solution. The deadline for submission is Friday, 9 October 2026, 23:59 BST.
We're not looking for standard solutions, we're looking for how you think. The strongest submissions are creative, original, and push beyond the obvious.
The goal of the challenge is to use real data collected by an applicant to drive a robotic manipulator in a simple simulation environment (e.g. Libero). The applicant is welcome to use a simple phone to record a small manipulation dataset and use it creatively showcasing their knowledge with VLA and/or World Models.
Here is an example of how it might look like:

Left: a snapshot from an egocentric hand-manipulation video sample; Right: simulation environment where we use that data to drive a Panda arm in Libero simulator with a trained policy based on SmolVLA.
Some suggestions on how you can develop your project:
use your recorded egocentric data to post-train a policy on a simple task
explore creative retargeting strategies, e.g. adapt your data to challenging embodiments
bootstrap a policy and use any form of RL to get better performance
use world modelling to showcase video/state prediction, less focusing on policy performance
optimise a standard policy to run considerably faster than a baseline
However, we don’t want to limit your imagination: in the age of AI agents, a standard task can now be easily achieved. You are welcome to use any resources at your disposal as long as the main constraint is achieved: you use the data that you personally collected. At the same time, we tried to design the challenge so that it could be hand-coded as well. We also checked that many ideas do not require access to large compute and could be done on Google Colab GPU notebooks.
Complete the challenge above and submit your solution as a public GitHub repository by Friday, 9 October 2026, 23:59 BST. Include a README/Presentation with instructions to run your system, example outputs, and a note on your design choices, what worked and what didn’t.
Creativity in approach, while satisfying the constraint: the data you collected must play a role in the approach
Performance of your policy in simulation and/or quality of WM predictions
Implementation simplicity and clear presentation of results without AI slop
Make something you’re proud of!
Location & Eligibility
Listing Details
- Posted
- September 28, 2026
- First seen
- October 2, 2026
- Last seen
- October 4, 2026
Posting Health
- Days active
- 1
- Repost count
- 0
- Trust Level
- 42%
- Scored at
- October 4, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.