ResearchOctober 07, 2026
10
š±
Robots finally learn from their own history
Long-WAM scales context for real-time robot control via AR pretraining.
#Robotics#World Models#Autoregressive#Real-time#Manipulation

š„ What happened
Researchers from Nvidia, HKU and others dropped Long-WAM, a world-action model that scales robot video context from 0 to 19.2 seconds. On RoboCasa GR-1, success jumps from 63.3% to 78.7%.
š” Why it matters
The trick isn't longer history per se ā it's how the video backbone was pretrained. Autoregressive pretraining pays off massively; bidirectional pretraining shows zero net gain. On an RTX 5090, one action chunk including future-video latent prediction runs in just 107.4 ms.
ā” Our take
Pi0.5 and Fast-WAM fail all 20 trials at dynamic cup stacking; Long-WAM nails 95%. If you're building robot policies without autoregressive pretraining, you're burning compute for nothing.
The title, summary and analysis of this item were produced automatically by an AI system and have not been editorially reviewed. They may contain errors, bias or omissions ā when in doubt, read the linked original source.
