Revisiting Latent Chain-of-Thought: Disentangling Representation and Supervision
Abstract
Latent chain-of-thought (CoT) reasoning aims to accelerate inference by replacing explicit text tokens with continuous hidden states. Existing approaches implement this concept through diverse latent representations and training objectives, but the tight coupling of these elements obscures the impact of individual design choices. To address this, we propose **LatentGrid**, a unified framework that decouples representation constraints from supervision strategies. By systematically varying representation constraints and supervision strategies under a shared experimental setup, we isolate their individual and joint effects on task accuracy, rationale recoverability, distributional alignment, and optimization dynamics. Our analysis indicates that unconstrained continuous representations paired with direct rationale-derived supervision better preserve recoverable reasoning information, while curriculum learning facilitates the transition to implicit reasoning. Conversely, restricting latent states to vocabulary embedding mixtures generally degrades performance and limits representational capacity. Inspired by these findings, we further propose **A**daptive **C**urriculum **T**hinking (ACT), a novel training paradigm that integrates curriculum learning with direct and token-level supervision. This synthesis improves reasoning accuracy and enables adaptive generation lengths, suggesting that effective latent reasoning benefits from flexible representations, explicit intermediate targets, and a progressive shift from explicit text.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.