acceptodds
Under review as a conference paper at ICLR 2027

Beyond Modality Importance: Token Routing Attention for Multimodal LLM Quantization

Abstract

Post-training quantization (PTQ) has become a key technique for reducing the storage and inference costs of multimodal large language models (MLLMs). Existing modality-aware PTQ methods learn independent activation smoothing scales for text and vision modalities, yet still rely on modality labels as the unit of quantization decisions, overlooking fine-grained variations in activation magnitudes across tokens within the same modality. In this paper, we propose Token Routing Attention Quantization (TRAQ), a PTQ framework that treats modality-specific activation smoothing scales as anchors and performs token-level scale routing through compatibility-based attention. Our core insight is that the optimal activation smoothing scale depends not only on global modality activation distributions but also on token-specific activation magnitudes. To this end, TRAQ treats modality-specific activation smoothing scales as prior anchors and employs a token-level attention mechanism that evaluates the compatibility between each token and each anchor, thereby assigning each token an activation smoothing scale aligned with its own activation magnitudes. On this basis, TRAQ computes layer-wise text and vision reconstruction errors and converts them into dynamic weights for the modality-specific reconstruction losses. The resulting weighted objective assigns greater emphasis to the modality with larger quantization error and is minimized to update the scale anchors. Extensive experiments show that TRAQ consistently outperforms competing PTQ methods, with gains of up to 7.70 points, while introducing no additional learnable modules and maintaining comparable inference efficiency to existing modality-aware PTQ methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.