HWAM: Humanoid World
Action Model
With Joint State–Action Generation
Joint state-action generation for closing the gap between what a humanoid robot is asked to do and what it actually does.
1 LimX Dynamics · 2 Southern University of Science and Technology · 3 Zhejiang University
*Equal contribution · †Corresponding author
Humanoid robots are a promising platform for general-purpose manipulation. Recent Vision-Language-Action policies and World Action Models learn to generate actions from multimodal observations, but hierarchical control introduces an action-execution gap: the reference produced by a policy can differ from the motion realized by the robot. We propose HWAM, a Humanoid World Action Model with joint state-action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target. HWAM connects policy references, realized body motion, and visual outcomes through three conditional training paths. On three real-robot tasks with the LimX OLI humanoid, HWAM achieves the highest success rate among evaluated baselines without pretraining on additional robot data.

The high-level policy proposes a reference action; whole-body tracking turns it into realized motion and a measured post-execution state.
Joint state-action generation for grounded control.
HWAM couples a video expert and a state-action expert through path-dependent mixed attention. The shared generator predicts future visual outcomes alongside the action and state that produce them.

One joint target,
three ways to learn.
Forward dynamics
Predict future visual observations conditioned on both actions and the states they produce.
→Inverse dynamics
Use visual transitions to reconstruct the actions and body states that caused them.
→Policy
Jointly denoise state-action trajectories from current visual observations, language, and proprioception.
→
Strong across stationary
and mobile tasks.
HWAM is evaluated on the LimX OLI humanoid across two stationary manipulation tasks and one mobile whole-body task.

| Method | Robot-data pretraining | Candy Picking | Object Collection | Plush-Toy Picking |
|---|---|---|---|---|
| Cosmos3-Edge | Yes | 15.8 | 0.0 | 56.7 |
| GR00T N1.5 | Yes | 40.5 | 42.5 | 60.0 |
| π0.5 | Yes | 62.5 | 45.0 | 60.7 |
| Fast-WAM | No | 43.3 | 0.0 | 60.0 |
| DiT4DiT | No | 52.5 | 15.0 | — |
| HWAM (Ours) | No | 70.6 | 46.7 | 73.3 |
Cite HWAM
Coming soon. Citation details will be available with the paper release.