INTACT: Inducing Latent Reasoning on a Frozen Backbone
Abstract
Latent reasoning replaces natural-language chains of thought with compact continuous “soft-token” representations, reducing the number of intermediate tokens that must be decoded autoregressively. We use a small external encoder to induce latent reasoning in a frozen backbone model. The encoder iteratively maps the backbone’s hidden states to new latent blocks, constructing a latent reasoning prefix, after which the frozen model writes its remaining reasoning and its answer as text. Our method, INTACT, trains this encoder through distillation to reproduce the effect of written reasoning steps while staying close to the frozen model’s behavior, and then reinforces it on answer correctness. INTACT raises GSM8K accuracy from 75.8% to 79.3% on Llama-3.2-3B and from 77.6% to 79.7% on Qwen-2.5-3B, and cuts end-to-end latency per problem by 19%. Accuracy improves on all four backbones without updating any backbone parameter, and INTACT outperforms the evaluated latent-reasoning baselines at the 3B and 8B scales. A separate encoder trained on a multi-domain corpus improves the same frozen backbone on twelve held-out mathematics, science and commonsense benchmarks without task-specific adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.