Reward Interactions and Precision Govern Resource Memory Before the Task
Abstract
Adaptive computation can leave unfinished work whose cost depends on later choices. In language models with early exits, for example, a later deeper token may require missing representations from earlier tokens. When a resource state must be summarized once for objectives revealed later, the summary must preserve information before its intended use is known. We study this requirement in a deferred-cost model with exact budget safety and an irreversible first commitment. We prove matching upper and lower bounds showing how reward interactions and numerical precision jointly govern the worst-case message size. At fixed regret, interactions confined to small groups permit logarithmic memory growth with the horizon, whereas interactions spanning a constant fraction of the horizon can require linear memory. Limited precision caps this growth, so a large requirement needs both extensive interactions and fine cost resolution. The lower bounds already hold for constructed prediction tasks whose optimizer, given the resource state, checks only a fixed number of plans. Preparing information for an unknown objective thus has a complexity distinct from optimizing that objective. We also identify supplied public cost structure under which an exact private summary remains short as shared rate precision increases, while charging computation separately. Exact finite checks and encoding experiments examine the constructions and their computational limits. The results clarify when reusable resource interfaces must retain history and when earlier objective disclosure or prospective plan queries can avoid that information cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.