Is WAM Really Better Than VLA?
For most of this year, people in robot learning have been saying that vision-language-action models (VLAs) are on their way out. Jim Fan put it most directly at Sequoia's AI Ascent in April, when he said VLAs could "rest in peace" and that world action models (WAMs), robot policies built on video generation models, would take their place1.
The leaderboards seem to back this up, with a WAM at the top of both NVIDIA's RoboLab2 and RoboArena3. The usual explanation is that video models learn how the physical world moves by watching it, while VLMs mostly learn to describe images, so a policy built on a video model should act better.
To see whether this holds up once the comparison is fair, I checked more than 90 papers against their own tables, seven leaderboards, 56 talks and posts, and two surveys, including the one I co-authored4. My short answer is that video pretraining clearly helps, but WAMs have not yet been shown to beat VLAs.
Q1: Do the leaderboards show that WAMs are better?
No, they rank systems, not ideas. None of the boards controls how much data or compute went into each entry, and the two in Figure 1 differ in exactly that: on RoboLab-1202 every entry is fine-tuned on DROID but brings its own pretraining, while on RoboDojo5 every entry also trains on the same 3,500 trajectories.
On RoboLab, a WAM, FLUX 3 Action6, leads π0.5 by 15 points, but π0 and π0.5 share a 3.3B architecture and differ by 23, so recipe and data alone can explain a gap that size. On RoboDojo, where the data is fixed, four of the nine WAMs beat π0.5 and five do not, and of four more boards run at least partly by outsiders, WAMs lead only RoboArena3, with a model four times larger.
Size is one of three things that change when a VLA is swapped for a WAM. VLAs sit at 2 to 4B on small VLMs such as PaliGemma7, while WAMs inherit their video generator's size, such as Wan2.18 at 14B for DreamZero9. FLUX 3 spends more than 95% of its base model's compute on video prediction10, while π0.511 trains on web image-text and robot data, and compute differs the most, as the official setups in Figure 2 show.
A 3.3B VLA fine-tuned on one GPU and a 16B WAM whose reference run used 256 GPUs differ by 5× in parameters and about 100× in fine-tuning compute, so a score gap between them could come from either. Both are reasonable engineering choices, but they are not a controlled experiment.
How stable are the baseline numbers?
Not very. π0.5 scores 58.35 on RoboTwin when the organizers run it, 43 to 44 in Motus15 and 77 to 83 in LingBot-VA16, a spread of 33 to 40 points that is larger than most WAM margins in these papers.
Papers also copy baselines from each other. The same set of RoboTwin baseline numbers appears in at least eight papers, two papers that copy ABot-M0 disagree by five points (81.2 and 86.1 on clean scenes)17,18, and Fast-WAM's19 LIBERO-Plus score is 51.5 in the papers that copy it and 60.0 in LAWA's re-implementation20.
Older benchmarks have also stopped separating models. In Ai2's aggregation of 3,971 reported results, 216 of 525 models score at least 97% on LIBERO21, which is why this post leans on newer boards.
Q2: Does the WAM lead come from size and compute?
Not mainly, as far as anyone has measured, but nobody has matched compute. Only AtomEgo22 gives both families the same training budget and no study matches FLOPs, while on RoboDojo5, whose paper lists the fine-tuning budget of 25 policies, a more than 80× spread in budget has a rank correlation of about 0.07 with success.
DreamZero's9 jump from 21% at 5B to 50% at 14B is the strongest case for size. But at matched steps the 14B model uses about 2.8× the training FLOPs, the VLA baselines score exactly 0%, which points to a failed recipe, and the two backbones may not be one family. Figure 3 puts every size comparison inside one WAM design on one plot.
Only DreamZero and Cosmos314 rise steeply. OpenWAM's25 14B backbone beats its 5B one by only 1.40 points, so the authors chose 5B, and Efficient-WAM26 distills 5B into 0.8B with no loss in score at a fifth of the latency. Latency is easier to compare than training cost, since several papers time a WAM and π0.5 on the same GPU.
Fast-WAM19 runs within 6% of π0.5 on the same GPU in its own paper, and Faster-WAM28, Efficient-WAM26 and GigaWorld-Policy-0.529 match or beat π0.5. Public WAM checkpoints are 5 to 47 times slower, and LingBot-VA up to 83 times in one setting.
Are the 5B and 14B models the same family?
The paper names the 5B backbone Wan2.1-I2V-5B-480P8. I could not find a public checkpoint by that name. The public 5B Wan model, which DreamZero's own repo uses for its 5B training script, is Wan2.2-TI2V-5B31, with a different VAE and a different training run. If that is the model used, size and backbone quality are confounded.
Q3: With the same data, is a video backbone better?
Video pretraining clearly helps, but it has not yet been shown to beat a VLM's pretraining. Starting from a video model beats starting from scratch in every ablation, while swapping a video backbone for a VLM gives split results once size and compute are taken into account (Figure 5).
The video backbone wins seven of the ten swaps, but most of those wins come with a second difference: DreamZero's WAM is about four times larger than its VLAs, the two LDA-1B23 models train on different data, and FLARE's33 video model trains five times longer and wins by 0.2 points. Only DiT4DiT32 matches model size, and there the video backbone wins by 2 points on LIBERO and 15 on RoboCasa-GR1.
AtomEgo22 is the only study that also holds training compute fixed, giving each model 3×8 B200s for about four days, and it splits: the VLA wins in distribution by 13 points and the WAM out of distribution by 4.
Against training from scratch, video weights win in all eleven papers that test it35,36,37,38,39,40,41,15,42,26,43, but mostly after the video model has seen manipulation, and Genie Envisioner38 gets near-zero success from a general one such as LTX-Video44. Robot pretraining data moves results as much as the backbone: WALL-OSS, a VLA pretrained on robot data, beats both WAMs by 27 points on ManipArena45, and ZimaBlue46 goes from 36.1% to 77.8% by changing only its pretraining data.
Q4: Does predicting future frames help?
Yes, but what helps is the model's internal representation of the future, not the rendered pixels. It helps most out of distribution and when data is scarce.
A WAM uses the future twice: it learns to predict future video during training, and its action head can look at the predicted future when it acts. Removing the first hurts in 18 of 19 ablations from 16 papers19,47,42,48,49,50,51,36,52,53,54,41,55,56,40,37, most on hard tasks: without video co-training, Fast-WAM19 drops 4 points on LIBERO but 65 on real towel folding. Switching off the second gives a wider spread (Figure 6).
The gain is near zero on LIBERO and full-data RoboTwin and reaches 64 points on held-out suites, so it is largest out of distribution and with little data.
The future does not need to be rendered, though. In the Simple-WAM study58, one pass over pure-noise future tokens recovers almost the whole gain, and in What Matters61, corrupting the generated future changes actions by less than 1% while reversing its time order costs 24 to 32 points.
Q5: Do WAMs generalize better out of distribution?
It depends on the shift. WAMs hold up better on new scenes, objects and tasks, and VLAs when the instruction or the task composition changes. Across released systems, training data matters more than whether a model is a WAM.
Zhang et al.30 test released WAMs and VLAs under perturbations. WAMs lead on RoboTwin 2.0-Plus and π0.5 leads on LIBERO-Plus, but the same Fast-WAM19 scores 72.7 on the first, where it trained on randomized scenes, and 51.5 on the second, where it saw only clean ones. Papers that train the VLA on the same data, or test both on the same benchmark, show more. Figure 7 collects 25 such comparisons, grouped by the kind of shift.
Most rows on new scenes, objects and tasks sit right of zero, with DiT4DiT32, Vidar64, Video2Act62 and DreamZero9 ahead by 15 to 70 points, and every row on instructions, recovery and composition sits left of it, by 16 to 24 points. The exceptions follow the training data and the sensors: Fast-WAM, trained on clean scenes, trails π0.5 by 44 points on randomized RoboTwin63, and WAMs trail when only the wrist camera is available66.
Q6: So, is WAM better than VLA?
Not yet: video pretraining is a real advantage, but no study has shown a WAM beating a VLA once data and compute are matched. Even Jim Fan has since narrowed his claim to fast reactive control71. Here is what each comparison in this post holds fixed.
| Study | Data | Params | Steps | Training compute | FLOPs | Real robot | OOD test |
|---|---|---|---|---|---|---|---|
| AtomEgo22 | ● | ○ | ● | ● same GPU time | n/r | ● | ● |
| DreamZero vs π0.5, GR00T9 | ● | ○ | ● | ○ | ○ | ● | ● |
| DiT4DiT32 | ● | ● trainable | ● | n/r | n/r | ● | ● |
| mimic-video72 | ● | ○ | n/r | ○ | ○ | n/r | ○ |
| Dyna-2, §04.173 | ● | n/r | ● | n/r | n/r | ● | ○ |
| Vidar vs π0.564 | ◐ fine-tuning only | ○ | n/r | n/r | ○ | ● | ● |
| Cosmos Policy vs π0.541 | ● | ○ | ○ | ◐ wall-clock | ○ | ● | ● |
| RoboDojo5 | ● | ○ | ○ | ○ | ○ | ● | ● |
| OpenWAM, Fast-WAM vs StarVLA7425,19 | ○ | n/r | n/r | n/r | n/r | ○ | ● |
| Zhang et al.30 | ○ | ○ | ○ | ○ | ○ | ○ | ● |
My guess is that the advantage is mostly robustness to how a scene looks, and that it comes from the representation rather than a 14B generator: the best open policy on RoboLab is half that size, and a 0.8B distilled video expert matches its 5B teacher. Architecture may not be the main lever at all, since π0.5 beats π0 by 23 points with the same one.
The fine-tuning half of a fair comparison runs on open weights today: Cosmos3-Edge against Cosmos3-Nano14, DreamZero at 5B and 14B9, and π0.5 and GR00T N1.712,13 on the same DROID data, all at matched GPU-hours. These are the questions I would most like answered; if you have run a compute-matched comparison, published or not, I would like to hear about it.
- At matched fine-tuning GPU-hours, how much of the WAM lead over π0.5 survives?
- Once training FLOPs are matched, does WAM performance keep improving with size, or does it flatten after a few billion parameters?
- Can a VLA get the WAM's robustness to visual shifts by adding a training-only future-prediction loss, as World Tokens48 and JEPA-WAM52 suggest, without giving up its lead on language and composition?
Acknowledgments
- Robotics’ End Game, Jim Fan, Sequoia AI Ascent, April 30, 2026 (talk). ↩
- RoboLab-120 leaderboard, NVIDIA. Numbers as of September 28, 2026. ↩
- RoboArena leaderboard, pulled September 29, 2026. ↩
- World Model for Robot Learning: A Comprehensive Survey, April 2026. ↩
- RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies, July 2026; leaderboard, numbers as of September 28, 2026. ↩
- FLUX 3 Action, Black Forest Labs, September 2026; fine-tuning guide. ↩
- PaliGemma: A versatile 3B VLM for transfer, July 2024. ↩
- Wan: Open and Advanced Large-Scale Video Generative Models (Wan2.1), March 2025; code. ↩
- World Action Models are Zero-shot Policies (DreamZero), February 2026; code. ↩
- FLUX 3 x mimic, Black Forest Labs, July 23, 2026. ↩
- π0.5: a Vision-Language-Action Model with Open-World Generalization, April 2025. ↩
- openpi (π0, π0-FAST, π0.5), Physical Intelligence, February 2025. ↩
- Isaac GR00T (N1.5 to N1.7), NVIDIA, March 2025. ↩
- Cosmos3-Nano-Policy and Cosmos3-Edge-Policy, NVIDIA, May and July 2026; training framework. ↩
- Motus: A Unified Latent Action World Model, December 2025. ↩
- Causal World Modeling for Robot Control (LingBot-VA), January 2026. ↩
- Motubrain: An Advanced World Action Model for Robot Control, April 2026. ↩
- Riemann-1.0: An Embodied World Action Model for Physical AI, August 2026. ↩
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?, March 2026. ↩
- Latent Action as Intention Enables Efficient Future Imagination for World Action Models, August 2026. ↩
- VLA Evaluation Harness leaderboard, Ai2, data of August 10, 2026. ↩
- AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining, September 2026. ↩
- LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion, February 2026. ↩
- ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?, June 2026. ↩
- OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining, September 2026. ↩
- Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination, June 2026. ↩
- Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA, May 2026. ↩
- Faster-WAM: Do World Action Models Need Deep Action Modules?, August 2026. ↩
- GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch, July 2026. ↩
- Do World Action Models Generalize Better than VLAs? A Robustness Study, March 2026. ↩
- Wan2.2, Alibaba Wan team, July 2025; Wan2.2-TI2V-5B weights. ↩
- DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control, March 2026. ↩
- FLARE: Robot Learning with Implicit World Modeling, May 2025. ↩
- Latent evolving World Action Model, September 2026. ↩
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations (VPP), December 2024. ↩
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation, December 2023. ↩
- VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation, November 2024. ↩
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation, August 2025. ↩
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation, September 2025. ↩
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators, December 2025. ↩
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning, January 2026. ↩
- GigaWorld-Policy: An Efficient Action-Centered World–Action Model, March 2026. ↩
- Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control, September 2026. ↩
- LTX-Video: Realtime Video Latent Diffusion, December 2024. ↩
- ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation, March 2026. ↩
- ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training, August 2026. ↩
- Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets (UWM), April 2025. ↩
- World Tokens: Enhancing Embodied Policies with Training-Time World Modeling, August 2026. ↩
- RynnVLA-002: A Unified Vision-Language-Action and World Model, November 2025. ↩
- WorldVLA: Towards Autoregressive Action World Model, June 2025. ↩
- InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization, July 2026. ↩
- JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling, August 2026. ↩
- BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation, February 2026. ↩
- The Right Future for Action: Learning Action-Relevant Predictive States in World Action Models, September 2026. ↩
- HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning, February 2026. ↩
- UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, January 2025. ↩
- Flex-π: A Multi-Stream World-Action Model with Compute Flexibility, August 2026. ↩
- What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling, September 2026. ↩
- Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models, August 2026. ↩
- GlanceWAM: Sparse Test-Time Imagination for World-Action Models, August 2026. ↩
- What Matters in Designing World Action Models: An Empirical Study, September 2026. ↩
- Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling, December 2025. ↩
- RoboTwin 2.0 leaderboard, updated September 24, 2026. ↩
- Vidar: Embodied Video Diffusion Model for Generalist Manipulation, July 2025. ↩
- RoboTwin-Phys: Do WAMs and VLAs Understand the Physical World?, September 2026. ↩
- LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models, September 2026. ↩
- RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents, September 2026. ↩
- Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio, July 2026. ↩
- LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models, September 2026. ↩
- Mask World Model: Predicting What Matters for Robust Robot Policy Learning, April 2026. ↩
- Panel: Robotics & World Models, Berkeley RDI, August 9, 2026 (video). ↩
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs, December 2025. ↩
- Dyna-2 technical report, Dyna Robotics, August 2026. ↩
- StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing, April 2026; code. ↩