acceptodds
Under review as a conference paper at ICLR 2027

Information-Rate Limits of Test-Time Reasoning

Abstract

Test-time reasoning systems spend compute by sampling candidate solutions and using a verifier to select among them, but more samples help only when the selection signal can identify the correct rollout. We formalize this coverage-precision gap by modeling verifier-guided inference as adaptive hypothesis testing through a feedback channel. The verifier's per-query information capacity yields an algorithm-independent lower bound on the number of queries needed for reliable recovery, with matching behavior in query-rich settings and collapse when the verifier is uninformative. Empirically, we hold rollouts fixed and vary only the selector across math, code, and multiple-choice benchmarks and two model families: oracle selectors expose latent gains, while practical selectors recover only part of them and corrupted verifiers collapse toward the single-sample baseline. A corrupted-oracle calibration places practical selectors—a trained process reward model, majority voting, an LLM-as-judge verifier, and log-probability ranking—on a common capacity-equivalent axis. Where information is directly measurable, the trained PRM and judge order consistently by realized information and captured gain, turning verifier quality into a measurable bottleneck for test-time compute.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.