acceptodds
Under review as a conference paper at ICLR 2027

Causal Inductive Biases for In-Context Learning in Diffusion Language Models

Abstract

Why does next-token pretraining produce in-context learners? Autoregressive (AR) training directly supervises prediction from the observed prefix, whereas masked and diffusion objectives distribute supervision across many observation patterns. We connect AR pretraining to Bayes-optimal episodic prediction and quantify the uncertainty reduction obtained by observing the future. These results motivate a supervision-allocation hypothesis: training more often on earlier context and later targets should promote sequential contextual adaptation. We test this hypothesis with matched AR, MLM and diffusion transformers, then introduce a minimal position- dependent corruption process. This intervention improves few-shot adaptation across diverse algorithmic tasks, survives native iterative generation under multiple diffusion objectives, and transfers to LLaDA-8B through continued pretraining. Together, these results identify training-objective factorisation as a central inductive bias for the emergence of in-context learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.