Why Does Non-Text Token Norm Inflation Emerge?
Abstract
Multimodal LLMs improve at understanding images and audio, yet a modality gap between non-text and text remains. We define as an increase in non-text token L2 norms during modality alignment, with these norms exceeding text token norms in early decoder layers. To examine its emergence, we track non-text norms during alignment. We then test its information-preserving role by replacing non-text tokens with inputs of varying norms but fixed directions. We assess the effects of mitigating norm inflation using during Qwen3.5-4B audio alignment. Our findings are threefold. First, early norm inflation is driven by increases in common-component norm. Second, we find an early decoder-layer range where replacing subsequent non-text tokens with inputs preserves at least 90% of originally correct predictions. Finally, TAN applied through early decoder layers yields lower accuracy after replacement than without TAN. TAN achieves higher inter-modal cosine similarity and improves performance across tasks, with a 9.9%p gain on MMAU. These findings support the information-preserving role of norm inflation. The results suggest that mitigating norm inflation may promote earlier formation of cross-modal representations and thereby improve task performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.