acceptodds
Under review as a conference paper at ICLR 2027

MemoSight: Improving the Fidelity of Chain-of-Thought Compression with Memory and Foresight Tokens

Abstract

Long chain-of-thought (CoT) reasoning enables LLMs to solve increasingly challenging tasks, but incurs substantial inference costs as the KV cache grows with the reasoning length. Online CoT compression reduces this footprint by replacing completed reasoning steps with compact memory tokens. However, existing methods compress every step into a *fixed, length-agnostic* set of memory tokens that are *positioned far from the content they summarize* and trained with *short-horizon next-token supervision*, limiting the fidelity of the compressed representation and degrading reasoning accuracy. We propose **MemoSight**, a lightweight framework that improves the fidelity of compressed reasoning while accelerating inference, through three designs targeting *distinct sources* of information loss. 1) **Adaptive Memory Allocation** allocates memory tokens in proportion to the length of each reasoning step, providing longer steps with greater capacity to preserve their information. 2) **Uniform Positional Layout** assigns memory-token position IDs uniformly within the span of the step they summarize, shortening their positional distance to the content and encouraging them to cover different parts of the step. 3) **Foresight Multi-Token Prediction** introduces foresight tokens that predict reasoning tokens farther ahead rather than only the next one, providing longer-horizon supervision that encourages richer memory representations. Across four reasoning benchmarks and three model families, MemoSight improves average reasoning accuracy by 2.6 points over LightThinker at comparable peak KV-cache usage. Moreover, the same foresight tokens used for training supervision are also used to generate draft predictions for speculative decoding, accelerating inference by 24.4%-49.4% across model families. Together, these results establish a better accuracy-efficiency trade-off for CoT compression.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.