Is WAM Really Better Than VLA?

For most of this year, people in robot learning have been saying that vision-language-action models (VLAs) are on their way out. Jim Fan put it most directly at Sequoia's AI Ascent in April, when he said VLAs could "rest in peace" and that world action models (WAMs), robot policies built on video generation models, would take their place1.

The leaderboards seem to back this up, with a WAM at the top of both NVIDIA's RoboLab2 and RoboArena3. The usual explanation is that video models learn how the physical world moves by watching it, while VLMs mostly learn to describe images, so a policy built on a video model should act better.

To see whether this holds up once the comparison is fair, I checked more than 90 papers against their own tables, seven leaderboards, 56 talks and posts, and two surveys, including the one I co-authored4. My short answer is that video pretraining clearly helps, but WAMs have not yet been shown to beat VLAs.

Q1: Do the leaderboards show that WAMs are better?

No, they rank systems, not ideas. None of the boards controls how much data or compute went into each entry, and the two in Figure 1 differ in exactly that: on RoboLab-1202 every entry is fine-tuned on DROID but brings its own pretraining, while on RoboDojo5 every entry also trains on the same 3,500 trajectories.

RoboLab-120 success rate versus parameter count Scatter plot of 14 policies on a log parameter axis with 95 percent confidence intervals. WAMs and VLAs are spread across the range; parameter count does not order the scores. RoboLab-120: each entry brings its own data 0 10 20 30 40 50 0.25B 1B 4B 16B parameters (log) success (%) FLUX 3 Action (WAM, 7B): 42.9% [95% CI 40.1–45.7] HiDream-O1-Embodied (VLA, 6B): 39.9% [95% CI 37.2–42.7], closed weights Atomic-WAM (WAM, 16.2B): 39.6% [95% CI 36.9–42.4] OASIS WAM (WAM, 16B): 39% [95% CI 36.3–41.8], closed weights Cosmos3-Nano-Policy (WAM, 16B): 36.8% [95% CI 34.1–39.5] Phoenix (other, 0.25B): 34.4% [95% CI 31.8–37.2], closed weights BiMind v0.1 (VLA, 7B): 33.3% [95% CI 30.7–36.1], closed weights π0.5 (VLA, 3.3B): 28% [95% CI 25.5–30.6] DreamZero (WAM, 14B): 25.7% [95% CI 23.3–28.2] Cosmos3-Edge-Policy (WAM, 4B): 22.9% [95% CI 20.6–25.4] π0-FAST (VLA, 3B): 15.5% [95% CI 13.6–17.7] GR00T N1.6 (VLA, 3B): 7.2% [95% CI 5.9–8.9] π0 (VLA, 3.3B): 5% [95% CI 3.9–6.4] paligemma-binning (VLA, 3B): 3.4% [95% CI 2.5–4.6] FLUX 3 Action Cosmos3-Nano π0.5 DreamZero Cosmos3-Edge π0 WAM VLA Other Closed weights RoboDojo simulation success by entry Sixteen RoboDojo entries trained on the same data, with types as labelled on the site. An agent and a VLA lead; the best three WAMs follow at 24 to 26 percent; four of nine WAMs score above pi0.5 at 6.93 percent and five below it. RoboDojo: every entry trains on the same data 0 20 40 simulation success (%) PhysicalRSI (Agent + VLA): 31.38% PhysicalRSI 31.38 Simate-beta (VLA): 27.96% Simate-beta 27.96 VPP2-Preview (WAM): 25.62% VPP2-Preview 25.62 Liber-0 Preview (VLM + WAM): 25.52% Liber-0 Preview 25.52 Liber-0 Lite (WAM): 24.23% Liber-0 Lite 24.23 GPT-6-Astra (Agent): 22.48% GPT-6-Astra 22.48 DM0.5 (VLA): 19.34% DM0.5 19.34 GalaxeaVLA (G0.5) (VLA): 14.88% GalaxeaVLA (G0.5) 14.88 Xiaomi-Robotics-1 (VLA): 13.93% Xiaomi-Robotics-1 13.93 OpenWAM-α (WAM): 11.92% OpenWAM-α 11.92 π0.5 (VLA): 6.93% π0.5 6.93 X-WAM (WAM): 3.83% X-WAM 3.83 GigaWorld-Policy-0 (WAM): 3.27% GigaWorld-Policy-0 3.27 AHA-WAM (WAM): 2.39% AHA-WAM 2.39 Fast-WAM (WAM): 2.03% Fast-WAM 2.03 LDA-1B (WAM): 0.54% LDA-1B 0.54 WAM VLA Agent
Figure 1. WAMs lead RoboLab, where each entry brings its own data, but not RoboDojo, where the data is shared (RoboLab-1202 with 95% intervals, RoboDojo5 simulation; both as of September 28, 2026).

On RoboLab, a WAM, FLUX 3 Action6, leads π0.5 by 15 points, but π0 and π0.5 share a 3.3B architecture and differ by 23, so recipe and data alone can explain a gap that size. On RoboDojo, where the data is fixed, four of the nine WAMs beat π0.5 and five do not, and of four more boards run at least partly by outsiders, WAMs lead only RoboArena3, with a model four times larger.

Size is one of three things that change when a VLA is swapped for a WAM. VLAs sit at 2 to 4B on small VLMs such as PaliGemma7, while WAMs inherit their video generator's size, such as Wan2.18 at 14B for DreamZero9. FLUX 3 spends more than 95% of its base model's compute on video prediction10, while π0.511 trains on web image-text and robot data, and compute differs the most, as the official setups in Figure 2 show.

GPUs needed to fine-tune and run each model The two VLAs fine-tune on one GPU; the three WAMs need at least eight, and NVIDIA's reference run for Cosmos3-Nano-Policy used 256. At inference the VLAs and FLUX 3 Action use one GPU and the two largest WAMs use two. fine-tuning GPUs 1 10 100 inference GPUs 1 2 π0.5 (3.3B): fine-tuning on 1 GPU(s) minimum; inference on 1 GPU(s) π0.5 VLA, 3.3B 1 1 GR00T N1.7 (3B): fine-tuning on 1 GPU(s) minimum; inference on 1 GPU(s) GR00T N1.7 VLA, 3B 1 1 FLUX 3 Action (7B): fine-tuning on 8 GPU(s) minimum; inference on 1 GPU(s) FLUX 3 Action WAM, 7B 8 1 DreamZero (14B): fine-tuning on 8 GPU(s) minimum; inference on 2 GPU(s) DreamZero WAM, 14B 8 2 Cosmos3-Nano-Policy (16B): fine-tuning on 8 GPU(s) minimum, reference run 256 GPUs; inference on 2 GPU(s) Cosmos3-Nano-Policy WAM, 16B 8 2 reference run: 256 GPUs (log) GPUs
Figure 2. The WAMs need eight GPUs to fine-tune where the VLAs need one (minimum official setups12,13,6,9,14).

A 3.3B VLA fine-tuned on one GPU and a 16B WAM whose reference run used 256 GPUs differ by 5× in parameters and about 100× in fine-tuning compute, so a score gap between them could come from either. Both are reasonable engineering choices, but they are not a controlled experiment.

How stable are the baseline numbers?

Not very. π0.5 scores 58.35 on RoboTwin when the organizers run it, 43 to 44 in Motus15 and 77 to 83 in LingBot-VA16, a spread of 33 to 40 points that is larger than most WAM margins in these papers.

Papers also copy baselines from each other. The same set of RoboTwin baseline numbers appears in at least eight papers, two papers that copy ABot-M0 disagree by five points (81.2 and 86.1 on clean scenes)17,18, and Fast-WAM's19 LIBERO-Plus score is 51.5 in the papers that copy it and 60.0 in LAWA's re-implementation20.

Older benchmarks have also stopped separating models. In Ai2's aggregation of 3,971 reported results, 216 of 525 models score at least 97% on LIBERO21, which is why this post leans on newer boards.

Q2: Does the WAM lead come from size and compute?

Not mainly, as far as anyone has measured, but nobody has matched compute. Only AtomEgo22 gives both families the same training budget and no study matches FLOPs, while on RoboDojo5, whose paper lists the fine-tuning budget of 25 policies, a more than 80× spread in budget has a rank correlation of about 0.07 with success.

DreamZero's9 jump from 21% at 5B to 50% at 14B is the strongest case for size. But at matched steps the 14B model uses about 2.8× the training FLOPs, the VLA baselines score exactly 0%, which points to a failed recipe, and the two backbones may not be one family. Figure 3 puts every size comparison inside one WAM design on one plot.

Success against video backbone size inside six WAM designs One line per WAM design, each on its own benchmark. DreamZero rises from 21 to 50 between 5B and 14B and Cosmos3 from 22.9 to 36.8 between 4B and 16B; OpenWAM, ImageWAM and LDA-1B rise by 2 to 5 points; Efficient-WAM keeps its score when distilled from 5B to 0.8B. 0 50 100 0.5B 1B 4B 16B video backbone size (log) score DreamZero (AgiBot): 5B 21, 14B 50 DreamZero AgiBot: 5B 21 → 14B 50 Cosmos3 (RoboLab-120): 4B 22.9, 16B 36.8 Cosmos3 RoboLab-120: 4B 22.9 → 16B 36.8 LDA-1B (RoboCasa-GR1): 0.5B 50.7, 1B 55.4 LDA-1B RoboCasa-GR1: 0.5B 50.7 → 1B 55.4 ImageWAM (LIBERO-Plus): 4B 83.1, 9B 85.2 ImageWAM LIBERO-Plus: 4B 83.1 → 9B 85.2 OpenWAM (RoboTwin 2.0): 1.3B 90.14, 2B 91.64, 5B 92.39, 14B 93.79 OpenWAM RoboTwin 2.0: 1.3B 90.14 → 14B 93.79 Efficient-WAM (distilled): 0.8B 87, 5B 86.4 Efficient-WAM distilled: 5B 86.4 → 0.8B 87
Figure 3. Only DreamZero and Cosmos3 gain much from a bigger video backbone (six WAM designs9,14,23,24,25,26, each on its own benchmark).

Only DreamZero and Cosmos314 rise steeply. OpenWAM's25 14B backbone beats its 5B one by only 1.40 points, so the authors chose 5B, and Efficient-WAM26 distills 5B into 0.8B with no loss in score at a fifth of the latency. Latency is easier to compare than training cost, since several papers time a WAM and π0.5 on the same GPU.

WAM inference latency as a multiple of π0.5 Latency per action chunk for WAMs, divided by π0.5 on the same GPU in the same source. Engineered WAMs that skip video generation sit near 1x; public checkpoints run 5 to 47 times slower, and up to 83 times for LingBot-VA in one setting. latency relative to π0.5, same GPU π0.5 WAM 1× 10× 100× Faster-WAM, measured in Faster-WAM, 24 GB GPU: 66.5 ms vs 71.4 ms, 0.9× π0.5 Faster-WAM 0.9× Efficient-WAM, measured in Efficient-WAM, RTX 4090: 98 ms vs 113 ms, 0.9× π0.5 Efficient-WAM 0.9× GigaWorld-Policy-0.5, measured in GigaWorld-Policy-0.5, RTX 4090: 110 ms vs 110 ms, 1.0× π0.5 GigaWorld-Policy-0.5 1.0× Fast-WAM, measured in Fast-WAM, RTX 5090D: 190 ms vs 180 ms, 1.1× π0.5 Fast-WAM (RTX 5090D) 1.1× Fast-WAM, measured in GigaWorld-Policy-0.5, RTX 4090: 182 ms vs 110 ms, 1.7× π0.5 Fast-WAM (RTX 4090) 1.7× Fast-WAM, measured in Faster-WAM, 24 GB GPU: 211.7 ms vs 71.4 ms, 3.0× π0.5 Fast-WAM (24 GB GPU) 3.0× Fast-WAM + future video, measured in Fast-WAM, RTX 5090D: 580 to 810 ms vs 180 ms, 3.2–4.5× π0.5 Fast-WAM + future video 3.2–4.5× GE-Act, measured in Zhang et al.: 4.8x, 4.8× π0.5 GE-Act 4.8× Cosmos Policy, measured in Zhang et al.: 6.2x, 6.2× π0.5 Cosmos Policy (Zhang et al.) 6.2× Cosmos Policy, measured in Beyond Task Success, 6000 Ada: 956 ms vs 100 ms, 9.6× π0.5 Cosmos Policy (6000 Ada) 9.6× Fast-WAM, measured in Beyond Task Success, 6000 Ada: 1418 ms vs 100 ms, 14.2× π0.5 Fast-WAM (6000 Ada) 14.2× Motus, measured in Zhang et al.: 18.6x, 18.6× π0.5 Motus (Zhang et al.) 18.6× Motus, measured in Efficient-WAM, RTX 4090: 3215 ms vs 113 ms, 28.5× π0.5 Motus (RTX 4090) 28.5× LingBot-VA, measured in Beyond Task Success, 6000 Ada: 4701 ms vs 100 ms, 47.0× π0.5 LingBot-VA (6000 Ada) 47.0× LingBot-VA, measured in Zhang et al.: 7.6x to 83x, 7.6–83× π0.5 LingBot-VA (Zhang et al.) 7.6–83× WAM latency ÷ π0.5 latency (log)
Figure 4. WAMs that skip video generation run about as fast as π0.5, while public checkpoints run 5 to 47 times slower (same-GPU timings from six sources27,28,19,26,29,30).

Fast-WAM19 runs within 6% of π0.5 on the same GPU in its own paper, and Faster-WAM28, Efficient-WAM26 and GigaWorld-Policy-0.529 match or beat π0.5. Public WAM checkpoints are 5 to 47 times slower, and LingBot-VA up to 83 times in one setting.

Are the 5B and 14B models the same family?

The paper names the 5B backbone Wan2.1-I2V-5B-480P8. I could not find a public checkpoint by that name. The public 5B Wan model, which DreamZero's own repo uses for its 5B training script, is Wan2.2-TI2V-5B31, with a different VAE and a different training run. If that is the model used, size and backbone quality are confounded.

Q3: With the same data, is a video backbone better?

Video pretraining clearly helps, but it has not yet been shown to beat a VLM's pretraining. Starting from a video model beats starting from scratch in every ablation, while swapping a video backbone for a VLM gives split results once size and compute are taken into account (Figure 5).

Success with a video backbone and with a VLM backbone Ten comparisons from six studies that swap a video backbone for a VLM backbone. The video backbone is ahead in seven, the VLM backbone in three, including AtomEgo in distribution. Only DiT4DiT matches model size and only AtomEgo matches training compute; the LDA-1B rows use different data, and FLARE's video model trains five times longer. video vs VLM backbone video backbone VLM backbone 0 50 100 DreamZero, AgiBot, new environments: video backbone 62.2, VLM backbone 27.4 DreamZero, new scenes 62.2 vs 27.4 DreamZero, AgiBot, unseen tasks: video backbone 39.5, VLM backbone 16.3 DreamZero, unseen tasks 39.5 vs 16.3 DiT4DiT, RoboCasa-GR1: video backbone 50.8, VLM backbone 36.2 DiT4DiT, RoboCasa-GR1 50.8 vs 36.2 AtomEgo, real, out of distribution: video backbone 27.14, VLM backbone 22.86 AtomEgo, real, OOD 27.14 vs 22.86 LDA-1B, RoboCasa-GR1: video backbone 55.4, VLM backbone 51.3 LDA-1B 55.4 vs 51.3 DiT4DiT, LIBERO: video backbone 98.6, VLM backbone 96.6 DiT4DiT, LIBERO 98.6 vs 96.6 FLARE, RoboCasa: video backbone 60.8, VLM backbone 60.6 FLARE, RoboCasa 60.8 vs 60.6 AtomEgo, real, in distribution: video backbone 30, VLM backbone 42.86 AtomEgo, real, ID 30 vs 42.86 FLARE, GR1: video backbone 29.5, VLM backbone 45.1 FLARE, GR1 29.5 vs 45.1 LeWAM, RoboTwin 2.0, frozen latent: video backbone 10.81, VLM backbone 79.2 LeWAM, frozen latent 10.81 vs 79.2 score (0 to 100)
Figure 5. A video backbone beats a VLM backbone in seven of ten swaps, but only DiT4DiT matches model size and only AtomEgo matches compute (six studies22,32,9,23,33,34).

The video backbone wins seven of the ten swaps, but most of those wins come with a second difference: DreamZero's WAM is about four times larger than its VLAs, the two LDA-1B23 models train on different data, and FLARE's33 video model trains five times longer and wins by 0.2 points. Only DiT4DiT32 matches model size, and there the video backbone wins by 2 points on LIBERO and 15 on RoboCasa-GR1.

AtomEgo22 is the only study that also holds training compute fixed, giving each model 3×8 B200s for about four days, and it splits: the VLA wins in distribution by 13 points and the WAM out of distribution by 4.

Against training from scratch, video weights win in all eleven papers that test it35,36,37,38,39,40,41,15,42,26,43, but mostly after the video model has seen manipulation, and Genie Envisioner38 gets near-zero success from a general one such as LTX-Video44. Robot pretraining data moves results as much as the backbone: WALL-OSS, a VLA pretrained on robot data, beats both WAMs by 27 points on ManipArena45, and ZimaBlue46 goes from 36.1% to 77.8% by changing only its pretraining data.

Q4: Does predicting future frames help?

Yes, but what helps is the model's internal representation of the future, not the rendered pixels. It helps most out of distribution and when data is scarce.

A WAM uses the future twice: it learns to predict future video during training, and its action head can look at the predicted future when it acts. Removing the first hurts in 18 of 19 ablations from 16 papers19,47,42,48,49,50,51,36,52,53,54,41,55,56,40,37, most on hard tasks: without video co-training, Fast-WAM19 drops 4 points on LIBERO but 65 on real towel folding. Switching off the second gives a wider spread (Figure 6).

Success with and without future tokens at inference Sixteen results from eight studies that keep the backbone and training data and change only whether the action head sees future tokens. The gain is near zero on LIBERO and full-data RoboTwin, and 10 to 64 points on LIBERO-Plus, held-out suites and small-data settings. action head with and without future tokens with future without 0 50 100 Fast-WAM, real towel folding: 0.70 with future, 0.75 without Fast-WAM, towel folding 0.75 → 0.70 Fast-WAM, RoboTwin 2.0: 90.6 with future, 91.8 without Fast-WAM, RoboTwin 91.8 → 90.6 Flex-π, RoboTwin 2.0, full data: 94.6 with future, 94.6 without Flex-π, full data 94.6 → 94.6 Fast-WAM, LIBERO: 98.5 with future, 97.6 without Fast-WAM, LIBERO 97.6 → 98.5 Simple-WAM study, LIBERO: 97.75 with future, 96.85 without Simple-WAM, LIBERO 96.85 → 97.75 Motus, RoboTwin 2.0, randomized: 87.02 with future, 83.9 without Motus 83.9 → 87.02 LAWA, RoboCasa few-shot: 63.1 with future, 54.5 without LAWA, RoboCasa 54.5 → 63.1 Efficient-WAM, 5B, RoboTwin 2.0, clean: 86.4 with future, 76.9 without Efficient-WAM, 5B 76.9 → 86.4 GlanceWAM, RoboCasa: 71.5 with future, 61.6 without GlanceWAM 61.6 → 71.5 LAWA, LIBERO-Plus: 70.4 with future, 60 without LAWA, LIBERO-Plus 60 → 70.4 Flex-π, real, 5 tasks: 63 with future, 50 without Flex-π, real 50 → 63 Simple-WAM study, LIBERO-Plus: 67.72 with future, 53.75 without Simple-WAM, LIBERO-Plus 53.75 → 67.72 Flex-π, RoboTwin 2.0, 50 demos: 60.4 with future, 40.2 without Flex-π, 50 demos 40.2 → 60.4 Efficient-WAM, 0.8B, RoboTwin 2.0, clean: 86.7 with future, 65.8 without Efficient-WAM, 0.8B 65.8 → 86.7 Faster-WAM, LIBERO-Plus: 73.57 with future, 51 without Faster-WAM 51 → 73.57 Simple-WAM study, held-out suites: 69.9 with future, 5.9 without Simple-WAM, held-out 5.9 → 69.9 score (0 to 100)
Figure 6. Looking at future tokens at inference changes little in some settings and adds up to 64 points in others, most on held-out suites, LIBERO-Plus and small datasets (eight studies19,57,15,26,20,58,59,60).

The gain is near zero on LIBERO and full-data RoboTwin and reaches 64 points on held-out suites, so it is largest out of distribution and with little data.

The future does not need to be rendered, though. In the Simple-WAM study58, one pass over pure-noise future tokens recovers almost the whole gain, and in What Matters61, corrupting the generated future changes actions by less than 1% while reversing its time order costs 24 to 32 points.

Q5: Do WAMs generalize better out of distribution?

It depends on the shift. WAMs hold up better on new scenes, objects and tasks, and VLAs when the instruction or the task composition changes. Across released systems, training data matters more than whether a model is a WAM.

Zhang et al.30 test released WAMs and VLAs under perturbations. WAMs lead on RoboTwin 2.0-Plus and π0.5 leads on LIBERO-Plus, but the same Fast-WAM19 scores 72.7 on the first, where it trained on randomized scenes, and 51.5 on the second, where it saw only clean ones. Papers that train the VLA on the same data, or test both on the same benchmark, show more. Figure 7 collects 25 such comparisons, grouped by the kind of shift.

WAM score minus VLA score out of distribution, by kind of shift Twenty-five out-of-distribution comparisons grouped by the kind of shift. WAMs lead on most new scenes, objects, environments and unseen tasks, and under randomized physics; VLAs lead on instruction following, recovery and task composition, and when only the wrist camera is available. out-of-distribution gap, WAM minus VLA WAM ahead VLA ahead -50 0 +50 New scenes, objects and environments DiT4DiT, real G1, new flower category: WAM 70, VLA 0, difference +70.0 DiT4DiT, new flower type +70 DreamZero, AgiBot, new environments: WAM 62.2, VLA 27.4, difference +34.8 DreamZero, new scenes +35 Video2Act, new lighting: WAM 43, VLA 19.7, difference +23.3 Video2Act, lighting +23 DiT4DiT, unseen objects, drawer: WAM 54.5, VLA 32, difference +22.5 DiT4DiT, unseen objects +22 Video2Act, new background: WAM 32.7, VLA 17.7, difference +15.0 Video2Act, background +15 RoboDojo, randomized scenes: WAM 4.11, VLA 1.44, difference +2.7 RoboDojo +3 OpenWAM, RoboTwin randomized: WAM 1.9, VLA 3.2, difference -1.3 OpenWAM, RoboTwin -1 RoboTwin 2.0, official, randomized, trained on clean: WAM 1.9, VLA 46, difference -44.1 RoboTwin 2.0, official -44 Unseen tasks Vidar, real, unseen tasks: WAM 67.1, VLA 12.9, difference +54.2 Vidar +54 DreamZero, AgiBot, unseen tasks: WAM 39.5, VLA 16.3, difference +23.2 DreamZero, unseen tasks +23 Physical properties RoboTwin-Phys, Fast-WAM vs π0.5: WAM 44.24, VLA 31.6, difference +12.6 RoboTwin-Phys, Fast-WAM +13 RoboTwin-Phys, Motus vs π0.5: WAM 39.6, VLA 31.6, difference +8.0 RoboTwin-Phys, Motus +8 Camera views LIBERO-VPro, contradicting views, LingBot-VA: WAM 65.8, VLA 52.5, difference +13.3 LingBot-VA, contradicting views +13 LIBERO-VPro, contradicting views, Fast-WAM: WAM 63, VLA 52.5, difference +10.5 Fast-WAM, contradicting views +10 LIBERO-VPro, wrist only, LingBot-VA: WAM 40.62, VLA 43.19, difference -2.6 (VLA: the weakest of three) LingBot-VA, wrist only -3 LIBERO-VPro, wrist only, Fast-WAM: WAM 19.69, VLA 43.19, difference -23.5 (VLA: the weakest of three) Fast-WAM, wrist only -23 Instructions, recovery, composition RoboFollow, instruction following: WAM 39.3, VLA 55.7, difference -16.4 RoboFollow -16 Temporal Ratio, composition, Cosmos Policy: WAM 10.5, VLA 26.8, difference -16.3 Cosmos Policy, composition -16 Temporal Ratio, composition, Fast-WAM: WAM 9.4, VLA 26.8, difference -17.4 Fast-WAM, composition -17 Temporal Ratio, composition, DiT4DiT: WAM 6.4, VLA 26.8, difference -20.4 DiT4DiT, composition -20 LIBERO-RECOVER, recovery from failures: WAM 16.7, VLA 40.8, difference -24.1 LIBERO-RECOVER -24 Mixed shifts AtomEgo, spatial, appearance, physical: WAM 27.14, VLA 22.86, difference +4.3 AtomEgo +4 Cosmos Policy, ALOHA, out of distribution: WAM 89.3, VLA 92.5, difference -3.2 Cosmos Policy, ALOHA -3 MWM, real, out of distribution: WAM 12.5, VLA 19.2, difference -6.7 MWM -7 OpenWAM, LIBERO-Plus: WAM 51.5, VLA 74.1, difference -22.6 OpenWAM, LIBERO-Plus -23 WAM score − VLA score (points)
Figure 7. WAMs hold up better on new scenes and tasks, VLAs on instructions and composition (WAM minus VLA in 25 comparisons from 15 studies32,9,62,5,25,63,64,65,66,67,68,69,22,41,70).

Most rows on new scenes, objects and tasks sit right of zero, with DiT4DiT32, Vidar64, Video2Act62 and DreamZero9 ahead by 15 to 70 points, and every row on instructions, recovery and composition sits left of it, by 16 to 24 points. The exceptions follow the training data and the sensors: Fast-WAM, trained on clean scenes, trails π0.5 by 44 points on randomized RoboTwin63, and WAMs trail when only the wrist camera is available66.

Q6: So, is WAM better than VLA?

Not yet: video pretraining is a real advantage, but no study has shown a WAM beating a VLA once data and compute are matched. Even Jim Fan has since narrowed his claim to fast reactive control71. Here is what each comparison in this post holds fixed.

What each comparison holds fixed: ● yes, ◐ partly, ○ no, n/r not reported.
StudyDataParamsStepsTraining computeFLOPsReal robotOOD test
AtomEgo22●○●● same GPU timen/r●●
DreamZero vs π0.5, GR00T9●○●○○●●
DiT4DiT32●● trainable●n/rn/r●●
mimic-video72●○n/r○○n/r○
Dyna-2, §04.173●n/r●n/rn/r●○
Vidar vs π0.564◐ fine-tuning only○n/rn/r○●●
Cosmos Policy vs π0.541●○○◐ wall-clock○●●
RoboDojo5●○○○○●●
OpenWAM, Fast-WAM vs StarVLA7425,19○n/rn/rn/rn/r○●
Zhang et al.30○○○○○○●

My guess is that the advantage is mostly robustness to how a scene looks, and that it comes from the representation rather than a 14B generator: the best open policy on RoboLab is half that size, and a 0.8B distilled video expert matches its 5B teacher. Architecture may not be the main lever at all, since π0.5 beats π0 by 23 points with the same one.

The fine-tuning half of a fair comparison runs on open weights today: Cosmos3-Edge against Cosmos3-Nano14, DreamZero at 5B and 14B9, and π0.5 and GR00T N1.712,13 on the same DROID data, all at matched GPU-hours. These are the questions I would most like answered; if you have run a compute-matched comparison, published or not, I would like to hear about it.

Acknowledgments

  1. Robotics’ End Game, Jim Fan, Sequoia AI Ascent, April 30, 2026 (talk). ↩
  2. RoboLab-120 leaderboard, NVIDIA. Numbers as of September 28, 2026. ↩
  3. RoboArena leaderboard, pulled September 29, 2026. ↩
  4. World Model for Robot Learning: A Comprehensive Survey, April 2026. ↩
  5. RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies, July 2026; leaderboard, numbers as of September 28, 2026. ↩
  6. FLUX 3 Action, Black Forest Labs, September 2026; fine-tuning guide. ↩
  7. PaliGemma: A versatile 3B VLM for transfer, July 2024. ↩
  8. Wan: Open and Advanced Large-Scale Video Generative Models (Wan2.1), March 2025; code. ↩
  9. World Action Models are Zero-shot Policies (DreamZero), February 2026; code. ↩
  10. FLUX 3 x mimic, Black Forest Labs, July 23, 2026. ↩
  11. π0.5: a Vision-Language-Action Model with Open-World Generalization, April 2025. ↩
  12. openpi (π0, π0-FAST, π0.5), Physical Intelligence, February 2025. ↩
  13. Isaac GR00T (N1.5 to N1.7), NVIDIA, March 2025. ↩
  14. Cosmos3-Nano-Policy and Cosmos3-Edge-Policy, NVIDIA, May and July 2026; training framework. ↩
  15. Motus: A Unified Latent Action World Model, December 2025. ↩
  16. Causal World Modeling for Robot Control (LingBot-VA), January 2026. ↩
  17. Motubrain: An Advanced World Action Model for Robot Control, April 2026. ↩
  18. Riemann-1.0: An Embodied World Action Model for Physical AI, August 2026. ↩
  19. Fast-WAM: Do World Action Models Need Test-time Future Imagination?, March 2026. ↩
  20. Latent Action as Intention Enables Efficient Future Imagination for World Action Models, August 2026. ↩
  21. VLA Evaluation Harness leaderboard, Ai2, data of August 10, 2026. ↩
  22. AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining, September 2026. ↩
  23. LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion, February 2026. ↩
  24. ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?, June 2026. ↩
  25. OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining, September 2026. ↩
  26. Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination, June 2026. ↩
  27. Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA, May 2026. ↩
  28. Faster-WAM: Do World Action Models Need Deep Action Modules?, August 2026. ↩
  29. GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch, July 2026. ↩
  30. Do World Action Models Generalize Better than VLAs? A Robustness Study, March 2026. ↩
  31. Wan2.2, Alibaba Wan team, July 2025; Wan2.2-TI2V-5B weights. ↩
  32. DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control, March 2026. ↩
  33. FLARE: Robot Learning with Implicit World Modeling, May 2025. ↩
  34. Latent evolving World Action Model, September 2026. ↩
  35. Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations (VPP), December 2024. ↩
  36. Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation, December 2023. ↩
  37. VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation, November 2024. ↩
  38. Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation, August 2025. ↩
  39. RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation, September 2025. ↩
  40. VideoVLA: Video Generators Can Be Generalizable Robot Manipulators, December 2025. ↩
  41. Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning, January 2026. ↩
  42. GigaWorld-Policy: An Efficient Action-Centered World–Action Model, March 2026. ↩
  43. Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control, September 2026. ↩
  44. LTX-Video: Realtime Video Latent Diffusion, December 2024. ↩
  45. ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation, March 2026. ↩
  46. ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training, August 2026. ↩
  47. Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets (UWM), April 2025. ↩
  48. World Tokens: Enhancing Embodied Policies with Training-Time World Modeling, August 2026. ↩
  49. RynnVLA-002: A Unified Vision-Language-Action and World Model, November 2025. ↩
  50. WorldVLA: Towards Autoregressive Action World Model, June 2025. ↩
  51. InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization, July 2026. ↩
  52. JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling, August 2026. ↩
  53. BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation, February 2026. ↩
  54. The Right Future for Action: Learning Action-Relevant Predictive States in World Action Models, September 2026. ↩
  55. HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning, February 2026. ↩
  56. UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, January 2025. ↩
  57. Flex-π: A Multi-Stream World-Action Model with Compute Flexibility, August 2026. ↩
  58. What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling, September 2026. ↩
  59. Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models, August 2026. ↩
  60. GlanceWAM: Sparse Test-Time Imagination for World-Action Models, August 2026. ↩
  61. What Matters in Designing World Action Models: An Empirical Study, September 2026. ↩
  62. Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling, December 2025. ↩
  63. RoboTwin 2.0 leaderboard, updated September 24, 2026. ↩
  64. Vidar: Embodied Video Diffusion Model for Generalist Manipulation, July 2025. ↩
  65. RoboTwin-Phys: Do WAMs and VLAs Understand the Physical World?, September 2026. ↩
  66. LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models, September 2026. ↩
  67. RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents, September 2026. ↩
  68. Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio, July 2026. ↩
  69. LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models, September 2026. ↩
  70. Mask World Model: Predicting What Matters for Robust Robot Policy Learning, April 2026. ↩
  71. Panel: Robotics & World Models, Berkeley RDI, August 9, 2026 (video). ↩
  72. mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs, December 2025. ↩
  73. Dyna-2 technical report, Dyna Robotics, August 2026. ↩
  74. StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing, April 2026; code. ↩