acceptodds
Under review as a conference paper at ICLR 2027

Reason Before You Speak with Latent Diffusion

Abstract

The simple next-token prediction objective leads to the emergence of remarkable downstream language capabilities. However, recent work has uncovered surprising failure modes of this objective on simple algorithmic tasks. Autoregressive models struggle to plan ahead, maintain global coherence, and explore diverse solutions. These limitations stem from the myopic nature of next-token prediction. We demonstrate that latent diffusion can imbue autoregressive models with foresight. By jointly training next-token prediction with diffusion over a latent representation of the future, models can first develop a holistic plan before committing to discrete tokens. We evaluate on a suite of algorithmic tasks designed to stress-test planning, coherence, and creativity. Across all settings, latent diffusion substantially improves over autoregressive baselines. We validate our approach on a mathematical reasoning benchmark. Our results suggest that augmenting language models with latent planning capabilities offers a promising path toward improving coherence, creativity, and reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.