Robotics
WAM
2025ExperimentalPublished
Key
innovation
Unified training of world prediction and action generation in a single autoregressive transformer — the model jointly learns physical dynamics (future visual observations) and robot policy (action sequences), enabling richer embodied representations without separate world-model and policy networks.
Category
Robotics
Abstraction level
Pattern
Operation level
ModelRobot controlTraining
Use cases
Long-horizon robot task planningSim-to-real transfer for manipulationGeneralist robot policies across environmentsInteractive robot assistants with goal-directed behaviorResearch on unified perception-action-prediction architectures
How it works
The model jointly learns to predict future world states and to generate action sequences conditioned on goals. It maintains an internal latent representation of the world that is updated as actions are taken, enabling multi-step planning grounded in physical consequences.
Problem solved
Current robot AI systems either plan abstractly (no motor grounding) or react reactively (no long-horizon planning). World Action Models unify predictive world modeling with direct action generation in a single architecture.
Components
Visual tokenizer
Action head
Future-frame decoder
Language conditioning
Implementation
Reference implementations
Implementation pitfalls
Kolaps na łatwiejsze zadanieCritical
Słaba tokenizacja akcjiHigh
Wysoki koszt obliczeniowy treninguHigh
Sim-to-real gap w rolloutMedium
Evolution
Hyperparameters (configurable axes)
Action tokenization schemeHigh
Future prediction horizonHigh
Action vs video loss weightingCritical
Decoder architectureMedium
Pretraining data mixHigh
Execution paradigm
Primary mode
Dense
Activation pattern
All paths active
Parallelism
Parallelism level
Partially parallel
Scope
TrainingAcross tokens
Hardware requirements
Primary
Good fit