CARD: Dense Denoising for Causal Language Models
Abstract
Autoregressive language models use causal computation efficiently but generate one token at a time. Masked diffusion permits parallel generation, although through a different training interface. Applying denoising to a causal model creates an uneven-context problem: the same corruption can remove most of the useful history for one target and little for another. Causal Autoregressive Diffusion (CARD) addresses this problem while retaining the standard next-token interface. It predicts every clean token from a corrupted left prefix, using Soft Tail Masking to preserve evidence within the corrupted suffix and context-aware weights to account for the amount and location of missing history. In matched 1B-parameter experiments, CARD exceeds block diffusion (BD3LM) by 5.10 percentage points in eight-task accuracy and trains about three times as fast. Ablations attribute the gain to both mechanisms, and prediction-level analysis finds less disagreement across corrupted prefixes. The same checkpoint supports KV-cached sequential generation and confidence-based parallel decoding, with parallelism selected at inference time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.