Humanoid manipulation · World action models

HWAM: Humanoid World
Action Model

With Joint State–Action Generation

Joint state-action generation for closing the gap between what a humanoid robot is asked to do and what it actually does.

Yan Yang1,*Jikun Rong1,*Minzhao Zhu1Zheyi Zhao1Qirui Hu1Zihan Lan1Weixin Mao1Yinhao Li1Zhen Fu2Hua Chen3,†

1 LimX Dynamics · 2 Southern University of Science and Technology · 3 Zhejiang University
*Equal contribution · †Corresponding author

▤ Paper
Robot in actionExecution, observed.Key task segments · looping previews
Abstract

Humanoid robots are a promising platform for general-purpose manipulation. Recent Vision-Language-Action policies and World Action Models learn to generate actions from multimodal observations, but hierarchical control introduces an action-execution gap: the reference produced by a policy can differ from the motion realized by the robot. We propose HWAM, a Humanoid World Action Model with joint state-action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target. HWAM connects policy references, realized body motion, and visual outcomes through three conditional training paths. On three real-robot tasks with the LimX OLI humanoid, HWAM achieves the highest success rate among evaluated baselines without pretraining on additional robot data.

01 / Overview
HWAM hierarchical control interface from policy reference through whole-body tracking to realized state

The high-level policy proposes a reference action; whole-body tracking turns it into realized motion and a measured post-execution state.

Architecture

Joint state-action generation for grounded control.

HWAM couples a video expert and a state-action expert through path-dependent mixed attention. The shared generator predicts future visual outcomes alongside the action and state that produce them.

HWAM architecture with multi-view video, language, current state, video expert, state-action expert, and joint prediction
HWAM architecture overview: visual context and joint state-action trajectories interact through two experts.
Three training paths

One joint target,
three ways to learn.

01

Forward dynamics

Predict future visual observations conditioned on both actions and the states they produce.

→
02

Inverse dynamics

Use visual transitions to reconstruct the actions and body states that caused them.

→
03

Policy

Jointly denoise state-action trajectories from current visual observations, language, and proprioception.

→
HWAM Policy, Forward Dynamics Modeling, and Inverse Dynamics Modeling training paths
These paths connect policy references, execution states, and visual transitions in both directions.
Real-robot evaluation

Strong across stationary
and mobile tasks.

HWAM is evaluated on the LimX OLI humanoid across two stationary manipulation tasks and one mobile whole-body task.

70.6%Candy Picking
46.7%Object Collection
73.3%Plush-Toy Picking
OLI stationary and mobile manipulation sequences
OLI evaluation suiteStationary and mobile manipulation in the real world.
Main results · success rate (%)
MethodRobot-data
pretraining
Candy
Picking
Object
Collection
Plush-Toy
Picking
Cosmos3-EdgeYes15.80.056.7
GR00T N1.5Yes40.542.560.0
π0.5Yes62.545.060.7
Fast-WAMNo43.30.060.0
DiT4DiTNo52.515.0—
HWAM (Ours)No70.646.773.3
Citation

Cite HWAM

Coming soon. Citation details will be available with the paper release.