acceptodds
Under review as a conference paper at ICLR 2027

Beyond Magnitude: Massive Activations in Diffusion Multimodal Language Models

Abstract

Diffusion multimodal language models generate responses through repeated masked-token updates, but the computational role of their largest residual activations remains unclear. We investigate these massive values in LaViDa, Dream-VL, and MMaDA through controlled interventions on fixed subsets of five multimodal datasets. Large components occur across image, question, template, and response tokens, yet their functional effects differ across regions, tasks, and models. Peak suppression can disrupt generation, but restoring the original norm does not consistently rescue performance, and retaining the peak alone is insufficient. Amplitude scaling further reveals asymmetric sensitivity. Routing peak access to downstream consumers reveals different dependencies on attention inputs, residual carry, normalized feed-forward reading, and terminal exposure. We develop a conditional theoretical account linking coordinate-concentrated writes and residual propagation to peak formation, and directional perturbations to discrete denoising decisions. These findings characterize massive values through their interaction with surrounding representations and downstream computation, rather than treating magnitude as a direct measure of semantic importance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.