Gradient Invasion and Alignment Tax in Unified Autoregressive Multimodal Models
Abstract
Unified autoregressive multimodal models train a single transformer and a single softmax over a mixed vocabulary of text subwords and visual codes. Prior theory of modality competition was developed for late fusion networks with separate encoders, and it does not determine the share when the bottleneck is one shared partition function. In this paper, we establish an exact last layer decomposition of the squared hidden gradient into a diagonal term and a cross term. On trained models the cross term is negligible, and the diagonal term therefore carries the share. That term is a Rayleigh quotient and yields a gradient share law in the visual token fraction, the residual energies, and an effective embedding scale. The operator norm provides a scale bound on this Rayleigh factor, and that bound does not replace the factor. Mixing a second discrete code stream into the same softmax raises text cross entropy, which we identify as a foreign code tax. A second text stream at the same mix ratio leaves text cross entropy essentially unchanged, while real visual codes and unstructured foreign codes raise it. A paired class caption does not raise it, so the excess is occupancy of the partition function by a foreign code. End layer isolation and inverse entropy reweighting act on different coordinates of an invasion plane, and the two interventions therefore cannot substitute for one another. The same measurements on public unified checkpoints show that gradient share is present for a joint softmax, a dual head, and a continuous visual head, while the scale channel exists only for a discrete shared softmax.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.