What Predictive-Improvement Rewards Make Agents Collect
Abstract
Predictive-improvement rewards value an observation by how much it helps predict later measurements. An agent that chooses those measurements can spend its budget making earlier discoveries earn credit. We study what this incentive makes it collect. With enough independent noiseless bits and an even budget, every policy maximizing the complete observation-deletion reward queries each chosen bit twice and acquires half the attainable information. An exact frontier characterizes the tradeoff, and a matching formulation shows how it depends on scoring times. The preference can also redirect learned behavior: switching an effective Gaussian-field policy to deletion reward increases credit, reduces coverage and leaves 45% more reconstruction error than continuing task training. Fitting masked predictions to policy records adds another dependence. Retained actions reveal deleted outcomes, changing the reward even with exact predictions; refitting can cycle or sustain a self-consistent optimum with lower information. Finally, interventions on temperature records show how unavailable inputs and changing evaluation targets alter the task value of selected measurements. These results connect the choices used to assign predictive credit to the information a finite acquisition budget buys.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.