acceptodds
Under review as a conference paper at ICLR 2027

Fork-Ready World Models: Training Distilled World Models to Be Branchable

Abstract

Tree-structured imagination relies on continuations sampled from a shared prefix, but a shared noisy cache can correlate those continuations and reduce the benefit of best-of- selection. We characterize this effect in a linear-Gaussian model and measure it in a B distilled autoregressive video model built on Self-Forcing and Wan2.1. The shipped clean-cache recipe has zero cache-noise variance in our protocol, whereas increasing cache noise reduces the measured variance ratio from to . We introduce Fork-Invariance, a training objective that compares noised-cache rollouts with clean-cache references using a shared sampling tape. Its strongest variance-ratio result () is misleading: full-weight training damages decoded video and reduces sensitivity to the prefix. Capping the objective's gradient norm at of the distillation gradient instead yields a variance ratio of , retains prefix sensitivity, and shows no detected quality degradation under two frozen visual judges. At matched measured variance ratio, the retained best-of- search gain increases from to . These results distinguish cache-noise suppression from preservation of useful conditioning and motivate evaluating branch groups through variance, selection performance, decoded quality, and prefix sensitivity together. The gradient-matched result uses one training run, and judge scores alone do not establish perceptual equivalence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.