ThinkBridge: Eliciting Latent Reasoning from Large Reasoning Models
Abstract
Explicit chain-of-thought (CoT) helps large language models (LLMs) solve problems, but generating long CoTs token by token incurs substantial inference latency. Existing latent reasoning methods fine-tune LLMs to carry intermediate computation in continuous representations. Reasoning-focused post-training has produced large reasoning models (LRMs), a class of LLMs with strong explicit reasoning capabilities. In preliminary mathematical reasoning experiments, we examine conventional supervised fine-tuning and existing latent reasoning methods on LRMs and find marked accuracy drops in some settings relative to the unmodified models. We also find that the original LRMs provide complementary successful responses in thinking and non-thinking modes. These observations motivate ThinkBridge, which keeps the LRM frozen and trains a lightweight reasoner using the model's own successful responses. The reasoner constructs compact latent states that guide the LRM to respond without first generating an explicit CoT. Training combines three objectives: supervision on correct native responses supports accurate answers; on-policy distillation aligns latent-conditioned and full-CoT-conditioned response predictions, training latent states to guide responses in place of CoT; and contrastive learning encourages question-specific latent states that support reasoning about the current question. We train ThinkBridge on GSM8K-Aug-NL. Across five mathematical benchmarks, it achieves the highest mean accuracy among the evaluated latent reasoning methods at 0.6B and 4B, with mean times to first response token (TTFT) of 45.50 and 56.33 ms, respectively. It also outperforms these baselines in multi-turn accuracy without dialogue-specific training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.