Self-prompted Latent Reasoners
Abstract
Large Language Models (LLMs) reason effectively when intermediate steps are written out as chain-of-thought (CoT) text, and latent reasoning models further move the explicit, token-level processes into internal computations in latent spaces before explicit output. For frozen LLMs, latent input tuning methods train a lightweight module that inserts continuous vectors after the task question, which improves the model's task-specific understanding and stability while avoiding catastrophic forgetting. However, the vectors produced by existing methods carry mostly an adaptation to the task rather than question-specific information, as replacing them with their average over questions leaves accuracy essentially unchanged, and these methods consequently perform poorly on tasks that require composing several reasoning steps. To address this issue, we propose self-prompted latent reasoners, in which a minimal mapper transforms internal representations of the intermediate layer of a frozen LLM and feeds them back to its input embeddings. The resulting signal is input-dependent and helps the model to read its own unfinished, richer thoughts, while it requires no auxiliary model and no additional input positions. Experiments on multi-step reasoning benchmarks show that the proposed method substantially outperforms prompt tuning and existing latent reasoning methods under a matched protocol, while remaining on par with them on common arithmetic benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.