acceptodds
Under review as a conference paper at ICLR 2027

Sentence Curve Language Models

Abstract

Language models (LMs) are central to modern AI systems, and diffusion language models (DLMs) have recently emerged as a competitive alternative. Both use word embeddings not only as inputs but also, implicitly, as target representations. We argue that conventional targets primarily provide token-wise supervision and do not explicitly encourage coordinated prediction across neighboring positions, which may be particularly important under noisy parallel decoding. To address this, we propose sentence curve, a spline representation whose control points are shared across multiple target words, and introduce the sentence curve language model (SCLM), which extends DLMs to predict through this shared representation. Theoretically, we show that sentence curve prediction reweights token-wise gradients and characterize a trade-off between target-side coupling and representation capacity. Empirically, SCLM achieves state-of-the-art performance among DLMs on IWSLT14 and WMT14, surpasses autoregressive Transformers on WMT14, and trains effectively without knowledge distillation. On LM1B, SCLM further improves generic language generation in a standard-scale fully non-autoregressive setting and shows promising preliminary transfer to semi-autoregressive block diffusion.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.