Diffusion-forcing world model rollouts. Each sample is ~2.7 s (16 raw frames @ 6 fps). Side-by-side videos show ground truth on the left, model prediction on the right. First half of each clip = history (GT), second half = autoregressively sampled future. Click a sample header to expand its videos (default-collapsed to keep the page light).
Click a run name to jump to its rollout videos. ✓ = feature on, ✗ = off. Lower val_loss is better.
| run | fusion | views | tactile | shift16 | delta-ref | cam-pose | gate | val_loss |
|---|---|---|---|---|---|---|---|---|
| p01rand_val_0.0140 | — | — | — | — | — | — | — | — |
Latent-space MSE of predicted future vs ground truth. Contact/no-contact split uses tactile latent-to-reference energy (threshold 0.05).
| run | view MSE | TL MSE | TR MSE | tactile contact MSE | tactile no-contact MSE |
|---|---|---|---|---|---|
| p01rand_val_0.0140 | 1.601581 | 1.981700 | 2.086621 | 2.038432 | 1.695410 |