acceptodds
Under review as a conference paper at ICLR 2027

How Training Objectives Allocate Predictive Information

Abstract

A representation with a fixed number of bits or a fixed linear rank must allocate information among the predictions it will support. We study how contrastive matching makes this choice when the prediction context arrives after encoding. A finite candidate list caps the gain from any one request. Matching can therefore favor information for harder requests even when prediction values it less. For two independent uniform bit strings with later complementary information, we solve prediction over all randomized encoders within one string's bit length and bound unavoidable predictive regret from attained matching scores. In two of six controlled neural settings, matching training loses information from a supplied prediction-optimal code. Simultaneous risk bounds at confidence for the fitted models show that a smoothed direct predictor outperforms every predictor of the paired matching code. Withholding complementary information reverses the minimum-risk ordering for the same parity-initialized pairs. We then characterize every optimal linear matching subspace in a Gaussian block model and quantify how predictive cost depends on optimization accuracy. Uncapped power divergences can favor a less frequent, more reliably observed string. Bounds over all randomized fixed-bit encoders distinguish bounded from growing regret as divergence order approaches one and string length grows. These results separate objective-induced information loss from errors in fitting a predictor.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.