acceptodds
Under review as a conference paper at ICLR 2027

Memorization Dynamics in Masked Diffusion Language Models

Abstract

Recent work on continuous diffusion models suggests that generalization and memorization can emerge on distinct training timescales, with population-level structure learned before individual training examples are memorized. In parallel, masked diffusion language models (DLMs) have become competitive language models while replacing autoregressive generation with iterative denoising. This raises a natural question: *does population learning also precede memorization in masked DLMs, or can the ordering change with the training regime?* We study this question using controlled Markov and hierarchical-grammar populations, trained canaries, never-trained controls, and Bayes-referenced population denoising. Contrary to a universal generalization-before-memorization picture, the relative timing of population learning and training-record extraction depends on the training regime. At low repetition, population learning can precede training-record extraction by thousands of optimization steps, whereas increasing repetition reverses the ordering and causes extraction to emerge first; in a contrasting Markov regime, extraction appears without an observed population-learning phase over the tested horizon. Across joint dataset-size and repetition interventions, normalized record frequency strongly organizes these dynamics, but controlled decoupling experiments show that it is not sufficient alone: both the sampling frequency of an individual repeated record and the aggregate training mass assigned to repeated records can alter the ordering. Mask-rate distribution shape further changes extraction onset even when its first two moments are matched, while pretrained LLaDA and MDLM reproduce the directional effect of repetition. Together, our results show that population learning and train-specific extraction are distinct phenomena whose relative timing is regime-dependent rather than universally ordered.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.