Causal Inductive Biases for In-Context Learning in Diffusion Language Models
Abstract
Why does next-token pretraining produce in-context learners? Autoregressive (AR) training directly supervises prediction from the observed prefix, whereas masked and diffusion objectives distribute supervision across many observation patterns. We connect AR pretraining to Bayes-optimal episodic prediction and quantify the uncertainty reduction obtained by observing the future. These results motivate a supervision-allocation hypothesis: training more often on earlier context and later targets should promote sequential contextual adaptation. We test this hypothesis with matched AR, MLM and diffusion transformers, then introduce a minimal position- dependent corruption process. This intervention improves few-shot adaptation across diverse algorithmic tasks, survives native iterative generation under multiple diffusion objectives, and transfers to LLaDA-8B through continued pretraining. Together, these results identify training-objective factorisation as a central inductive bias for the emergence of in-context learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.