Learning to Correct Is Not Enough: Reusing Feedback in World Models
Abstract
Accurate short-path correction does not guarantee reliable reuse along longer action paths. We study how a single observation can revise multiple cached futures of a world model without regenerating them. A matched construction adds one known return without changing the unknown response or calibration data, turning horizon-independent minimax decision risk into risk growing with path length until saturation. Legal continuations determine which response uncertainty is repeatedly exposed to the task. For known compact classes of orthogonal response dynamics on a fixed finite directed graph, we characterize when Gaussian-noisy scalar observations on a prescribed short-path bank suffice for reliable decisions along all legal continuations. Learning is possible exactly when, in every strongly connected component and at every positive task precision, only finitely many irreducible types of the joint return action remain task-significant across the candidate class. This permits infinitely many distinct return types when their task influence vanishes. Compatible internal coordinate changes cancel in composition, bounding approximation error per visited component rather than per action; under this criterion, longer paths need not demand finer calibration. We further characterize when all path responses admit a fixed finite exact linear calibration with model-independent coefficients, while showing that such exact linear tables can require exponentially many entries. On a continuous nominal base, even a smooth, uniformly observable class sharing one known response type can preclude uniform all-path learning from noisy short-path data. Controlled studies and simulated planning examine these distinctions and the trade-offs of cached correction. Calibration sufficiency therefore depends not only on what the data reveal, but also on which unresolved response differences future use can make consequential.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.