acceptodds
Under review as a conference paper at ICLR 2027

Why Does Non-Text Token Norm Inflation Emerge?

Abstract

Multimodal LLMs improve at understanding images and audio, yet a modality gap between non-text and text remains. We define as an increase in non-text token L2 norms during modality alignment, with these norms exceeding text token norms in early decoder layers. To examine its emergence, we track non-text norms during alignment. We then test its information-preserving role by replacing non-text tokens with inputs of varying norms but fixed directions. We assess the effects of mitigating norm inflation using during Qwen3.5-4B audio alignment. Our findings are threefold. First, early norm inflation is driven by increases in common-component norm. Second, we find an early decoder-layer range where replacing subsequent non-text tokens with inputs preserves at least 90% of originally correct predictions. Finally, TAN applied through early decoder layers yields lower accuracy after replacement than without TAN. TAN achieves higher inter-modal cosine similarity and improves performance across tasks, with a 9.9%p gain on MMAU. These findings support the information-preserving role of norm inflation. The results suggest that mitigating norm inflation may promote earlier formation of cross-modal representations and thereby improve task performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.