acceptodds
Under review as a conference paper at ICLR 2027

CLUE: Contrastive Learning to Unlock Effective Latent Reasoning in LLMs

Abstract

Latent reasoning has emerged as a popular alternative to textual chain-of-thought by carrying intermediate information through continuous representations. As LLMs are pretrained only on text rather than continuous states, post-hoc training is required for them to effectively process latent representations. However, under standard outcome-based reinforcement learning, this inherent text bias can limit the model’s effective use of latent information, as the objective rewards final-answer correctness without explicitly encouraging the model to exploit the information conveyed by the latent representations. We address this challenge by introducing a decoupled architecture paired with a novel training objective. Because latent thoughts represent a fundamentally different knowledge formulation than text, our architecture separates latent generation from textual decoding: a trainable latent generator guides a lightweight, frozen text model, forcing learned adaptation exclusively through the latent channel. To explicitly encourage reliance on these continuous thoughts, we introduce a Contrastive Learning objective to Unlock Effective latent reasoning in LLMs (CLUE). CLUE adds a contrastive loss to the standard GRPO objective, ensuring that correct responses yields a higher log-likelihood under their own positive latent compared to negative latents sampled from other inputs. This provides dense supervision for latent utilization alongside answer correctness, without requiring ground-truth reasoning traces. Trained on OpenThoughts30K and evaluated across diverse mathematical reasoning and scientific question-answering benchmarks, our method improves average accuracy by 5.6 percentage points over direct optimization of the text-based model and outperforms related work by up to 3.5 percentage points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.