What Makes a Good Latent Thought for Reasoning?
Abstract
Latent reasoning promises shorter generation by replacing token-level traces with compact continuous thoughts, but these representations must both preserve the reasoning trace and be reliably generated from the input. We ask which properties of latent thoughts support reliable reasoning under compression, using controlled permutation state-tracking tasks and varying compression and representation structure. Our findings are threefold. (i) At moderate compression, autoregressive (AR) generation remains highly accurate with far fewer emissions: 20 latent emissions recover a 160-token trace with near-perfect final-state accuracy, matching or exceeding a parameter-matched token-AR baseline with \(8\times\) fewer emissions. (ii) At higher compression, reconstruction remains accurate while generation fails. Factorization or blockwise encoding restores accuracy, showing that latent structure strongly affects the ability to generate. With further compression, errors persist even with correct prefixes and on a simpler affine composition task, revealing a composition-capacity limit on how many state transitions the generator can realize within one emission. (iii) Iterative refinement is not a universal remedy. Diffusion does not recover highly compressed composition. On context-dependent graph walking, however, compression reverses the generator ordering: AR falls from perfect to chance, while diffusion improves. Together, these results show that usable compression depends on both latent structure and the computation packed into each emission. Effective latent reasoning therefore requires the right computational granularity for the representation, generator, and task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.