acceptodds
Under review as a conference paper at ICLR 2027

Latent Reasoning Does Not Lack Signal, It Lacks Invariance: Subspace-Constrained Refinement for Robust Reasoning

Abstract

Latent reasoning methods in LLMs perform multi-step inference in embedding space, avoiding the token overhead of Chain-of-Thought while preserving reasoning capacity. Recent training-free refinements steer these latent trajectories using contrastive signals from reference checkpoints, yielding 2-5% accuracy gains at zero parameter cost. We show that these gains are fragile along three aspects: (i) hyperparameter sensitivity, where the step size and memory rate have task-specific optima with no per-instance adaptation; (ii) reference-pair dependence, where the contrastive direction from a single checkpoint pair can reflect pair-specific artifacts rather than stable reasoning structure; and (iii) input sensitivity, where updates in the full embedding space let surface features push equivalent inputs onto different trajectories. We argue that reliable latent refinement requires trajectory-level stability, not only answer-level accuracy, and that the refinement must degrade gracefully when its own signal is unreliable. We propose Trajectory-Invariant Latent Refinement (TILR), a training-free framework with two components. First, we identify a low-rank subspace in which reference models consistently differ across a calibration set, and project contrastive updates onto this subspace to filter unstable variation. Second, we apply an adaptive gating mechanism that scales update magnitude according to the alignment between the current signal and the learned subspace, suppressing unreliable corrections and recovering the unmodified backbone when alignment is low. Across six reasoning benchmarks, TILR consistently improves over prior refinement methods while substantially reducing sensitivity to input perturbations and reference selections. In particular, TILR improves answer agreement under semantically equivalent paraphrasing by 11.4% and reduces trajectory variance by  30% relative to the Coconut backbone, while maintaining comparable accuracy. These results demonstrate that constraining refinement to invariant subspaces yields more reliable latent reasoning without requiring retraining or backpropagation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.