Compute, Fuse, or Acquire? Test-Time Decision-Information Bottlenecks
Abstract
Test-time scaling is usually treated as allocating more computation, or more interaction with the world. We show that this conflates three distinct bottlenecks. Computational regret is imperfect inference over the current interface; representation regret is evidence the model has already observed but its interface discards; observation regret is evidence never observed. The algebra separating them is elementary, but the consequence is not: computational, representation and observation regret require distinct remedies, namely compute, fusion and acquisition respectively. This yields a prediction a compute-versus-interaction view cannot state, namely that performance can improve at exactly fixed evidence. We test the predictions on InfoGap, a benchmark whose floors are designed and, for vision, tunable. Across three VLMs, varying the designed single-view Bayes accuracy moves the observed sampling plateau with it, separating an information ceiling from a capability limit; fixed-view sampling stays at that ceiling while one relevant view lifts accuracy from 0.49 to 0.99. Holding the observed evidence identical and changing only how two channels are combined moves accuracy from 0.50 to 1.00, which acquisition cannot explain. A bottleneck-matching router is a proof of concept rather than a transferable diagnostic: it reaches 0.94 within a known task mixture at half the generated-token cost of Best-of-16, but leave-one-category-out shows it identifies task families, not instances.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.