acceptodds
Under review as a conference paper at ICLR 2027

LatentLM: Learning a Discrete Internal Language for Large Language Model

Abstract

In this paper, we study whether an off-the-shelf LLM can be adapted into a discrete, variable-length token compressor and decompressor for long-context processing. To this end, we design a self-expressive autoencoding framework that fine-tunes a pretrained LLM with lightweight LoRA adapters to map long texts into compact sequences of learned latent codes, termed Z-tokens, and to decode them back into natural language or task outputs. The resulting representation is content-adaptive: less predictable or more information-dense segments can receive more Z-tokens, while redundant regions can be represented more compactly through a budget-aware length regularizer. We evaluate our method on seven long-context datasets (e.g., RULER, HotpotQA, and QuALITY), demonstrating that, compared with several strong baselines, our approach achieves better reconstruction quality and downstream performance, especially in scenarios with multiple queries on the same input, while reducing effective context length and memory usage and improving throughput in the generation stage. This design supports direct decoding from compressed contexts, and we further explore promising paradigms for autoregressive generation in the Z-token space, providing a practical interface for efficient long-context reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.