Efficient Multimodal Contrastive Decoding via Budgeted Activation
Abstract
Multimodal contrastive decoding (MCD) has been widely adopted to mitigate hallucination in multimodal large language models (MLLMs). It is performed by contrasting predictions from an original input with those from an auxiliary branch. The inference latency thus increases to at least *2×* that of vanilla decoding. However, our preliminary experiments show that MCD sparsely alters the output token (e.g., 7% on POPE). This mismatch between computational cost and the rate of altered tokens motivates a token-level analysis of MCD's activation. Accordingly, we introduce **FlipTrace**, a theoretical framework that characterizes MCD's effect as *flips*. We then prove a *flip-immunity condition* that bounds how many tokens MCD can affect and show that these effects concentrate where the bound predicts. Together, these findings lead to **flip susceptibility**, a statistic computed from the original logits alone that predicts token flips. Using this measure, we develop a **training-free budgeted activation** method. Within this budget, our method prioritizes MCD activation on decoding steps with high flip susceptibility. We evaluate our method across diverse combinations of *six MCD methods, four MLLM backbones, and five benchmarks*, spanning visual and audio-visual tasks. Across these settings, budgeted activation integrates with existing MCD methods to achieve a maximum reduction of **87% in inference latency**, while preserving their performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.