Beckmann World Models

Direct Terminal-Map Learning for
One-Step Action-Conditioned Prediction

Adam LeeBerkeley Shobhit AgarwalStanford JP SuhGoogle DeepMind
Conditional Beckmann terminal-map conservation, source supervision, and motion and temporal losses
BWM combines terminal-map transport training with direct supervision of predictions from noise. Inference uses one call per RGB chunk or latent frame.

Qualitative results

PushTSpatial feature supervision

Recorded framesBWM

Robomimic CanSpatial feature supervision

Recorded framesBWM

Bridge-V2Population refinement · frames 17–44

RT-1Population refinement · frames 17–89

Selected examples. RGB: four observed and 60 predicted frames. Native: full-data population refinement, VAE-reconstructed references, recorded context reset every eight frames.

Matched PushT and Can predictions across ground truth, DriftWorld, regression, GPC, AVDC, and BWM with spatial supervision, at initial and extended fits

Quantitative results

RGB prediction

ModelPushTRobomimic Can
MSE ↓Motion MSE ↓LPIPS ↓MSE ↓Motion MSE ↓LPIPS ↓
MSE U-Net 0.03030.40080.08040.00510.05440.0104
DriftWorld 0.03140.40930.07900.01040.07690.0682
GPC 0.06301.42410.16090.02760.31030.0589
AVDC 0.06841.18400.18600.04000.26440.1194
BWM + spatial 0.00980.12470.02800.00220.02420.0056

388 PushT and 30 Can episodes. Four observed and 60 autoregressive predictions at 96 × 96 resolution.

Native video prediction

ModelBridge-V2RT-1Calls per frame
SSIM ↑PSNR ↑LPIPS ↓SSIM ↑PSNR ↑LPIPS ↓
Matched subset · 2,078 Bridge-V2 and 2,041 RT-1 training episodes
DriftWorld Population drifting objective.81321.919.104.81422.419.1291
MSE U-Net.81322.439.166.81523.026.1791
GPC.64219.932.335.69420.692.2943
AVDC.74619.595.159.74920.269.175100
BWM spatial .79622.130.155.80223.147.1661

215 Bridge-V2 and 198 RT-1 episodes. Recorded context resets every eight frames. Metrics use VAE-reconstructed references.

Can errors from pure-noise to data-boundary inputs
Can LPIPS versus cumulative training-loop hours including inherited training