LTC: Risk-Gated Lagrangian Token Compression for Video Large Language Models
Abstract
Video Large Language Models (video-LLMs) process large numbers of visual tokens, incurring substantial computational and memory overhead. Existing training-free compression methods reduce spatiotemporal redundancy through importance-based selection, cross-frame matching, and structured aggregation. However, these methods neglect the local motion structure of neighboring correspondences, which may lead to the loss of key dynamic information during compression. Based on this insight, we propose risk-gated Lagrangian Token Compression (LTC), a training-free inference framework for video-LLMs. Specifically, Token Flow Estimation (TFE) uses local feature matching to estimate token displacements and matching confidence. Eulerian Risk Diagnosis (ERD) then combines neighborhood motion structure with active low-confidence signals to construct an aggregation-risk proxy. Guided by this risk proxy, Risk-Gated Lagrangian Aggregation (RGLA) keeps protected tokens as independent candidates, aggregates the remaining tokens that satisfy feature-consistency constraints into compact representatives along temporal trajectories, and selects the final visual tokens by importance under a token budget. The framework directly reuses pretrained visual features without training an additional network. Experiments on three representative video-LLMs and three video understanding benchmarks validate LTC's effectiveness and applicability across models. Retaining only 25% of visual tokens, LTC achieves 101.3% of the vanilla model's average performance on LLaVA-OneVision. Code is available at https://anonymous.4open.science/r/LTC-75E4.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.