Pricing Presented Information: A Dual-Control Estimand for Reward Sacrifice in Language Models
Abstract
Actions in interactive tasks can trade immediate reward for feedback that improves future decisions. We give a behavioural estimand for how much a language-model policy pays for such feedback: the exchange rate between immediate success probability and expected feedback information that its choices reveal, recovered by a conditional logit with a neutralised negative control. Dual control theory motivates and delimits the estimand: an optimal policy must price probing, and probing has strict value exactly when feedback can push the posterior across a future decision boundary. We build benchmarks where this price has bite: 161 validated Mastermind and Wordle states in which guaranteed success requires a reward-sacrificing probe, and a structurally different fault-repair environment with noisy feedback and stochastic reward. Across ten systems from two open-weight families and three frontier models, the estimand is well identified, robust to decoding temperature, symbol identity and reasoning budget, and differs sharply between systems in ways parameter count does not explain. The central finding concerns what is being priced. Masking and deception tests show that the valuation attaches to legible presented statistics: scrambling them returns models to certainty-equivalent play, and hiding them collapses Sonnet 5's valuation in both environments and Qwen-32B's in fault repair, where the tables from which information value could be computed remain on screen. Consistently, the valuation is environment-specific: it predicts success in the puzzles (Spearman ), but in pre-registered tests neither the ranking of models nor the link to success transfers to fault repair. The price language models put on information is a measurable property of how they use presented statistics, not a portable trait.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.