acceptodds
Under review as a conference paper at ICLR 2027

The Correction Budget: When Test-Time Refinement Can and Cannot Fix World-Model Rollouts

Abstract

Test-time refinement—predict a latent, then project it onto a learned constraint—is an appealing recipe against the compounding drift of action-conditioned world models. We build the recipe in its strongest practical form (a transition energy model trained against the world model's own drifted rollouts, fixing verifier AUROC from 0.06 to 0.88–0.94) and then take it apart. Across three environments, two observation modalities, and fifteen (world-model, corrector) systems, we find: (i) drift does not compound exponentially—the off-manifold amplification premise fails; (ii) the verifier is a pure legality detector: its gradients, regression targets, and selection weights carry no information about the direction back to the truth; (iii) a maximal-input probe—the world model fine-tuned on its own rollout contexts—shows that open-loop error has a predictable fraction governed by two separable channels, reducible model bias (data starvation) and off-manifold geometry (representation); (iv) a correction-budget law: no test-time corrector measurable in the rollout's inputs can beat . The bound is saturated exactly where per-step alignment is high—pixel latents, where legality truth and refinement recovers up to 46% of drift—and correctly forecasts the near-zero gains on near-manifold proprioception.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.