acceptodds
Under review as a conference paper at ICLR 2027

Two Channels of Sequential Evidence: Why a Better Predictive Representation Did Not Detect Change Better

Abstract

A representation that predicts the future well ought to make a change in the world easier to notice. We set out to demonstrate exactly that, on a benchmark built so that only genuinely nonlinear structure can survive, with the detector held fixed so that every representation is judged on identical terms and with the interpretation of every outcome fixed before the numbers were seen. The result was the opposite of the one we expected. The learned predictive representation recovered the underlying state better than anything else we trained, and still detected change no better than a linear method from the subspace-identification literature that costs a single matrix factorisation and no gradient steps at all; under a second, equally legitimate detector the ordering reversed outright. Chasing that failure proved more informative than the success would have been. We show that the evidence a sequential test accumulates splits, exactly and by inspection of the test itself rather than by modelling it, into two separate channels: one that keeps paying as evidence accrues, and one that pays a bounded amount and then stops. Whether a representation helps is a question of which channel it opens, and every representation we trained left the paying channel shut while the true state opened it. From this we prove that a representation cannot manufacture evidence at all: it can only fail to destroy it, or move it into a form the test is able to read. That turns two of our most damaging measurements from embarrassments into predictions. We further find that the ordering of representations is a property of the extreme tail of the test statistic rather than of any average, so an entire family of cheap diagnostics, including one we pre-registered and one we had to retract mid-study, could never have worked. We report the negative result, the machine-checked theory that explains it, and a measurement that turns “which channel does this benchmark open” from an assumption into an answerable question.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.