Compute Once, Update Many: Sharing Computation for Efficient Video-LLM Understanding
Abstract
Video-language models spend substantial computation repeatedly processing visual tokens during language decoder prefill. Existing acceleration methods often reduce this cost by removing or merging tokens, but retaining a token's individual representation need not require computing its update independently. We introduce a carrier–receiver computation sharing method that reduces repeated computation while updating each visual token's own hidden state. Selected visual carriers compute Attention and feed-forward network (FFN) updates, which receivers add to their own residual states. All token positions and key/value entries are retained, and non-visual tokens are processed normally. Input-dependent assignments identify which tokens share updates, while calibrated layer-wise ratios control the amount of sharing. Across three model families, our method maintains accuracy close to dense inference on MVBench, VideoMME, and LongVideoBench, while delivering measured decoder prefill speedups of – at 32 frames and consistent acceleration from 8 to 64 frames. Beyond standalone acceleration, our method consistently outperforms further token reduction alone when combined with four visual token reduction methods, improving accuracy by up to percentage points at matched decoder prefill FLOPs across the three benchmarks. These results show that sharing updates offers an effective way to reduce decoder computation while keeping visual tokens available for subsequent attention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.