CueMiner: Efficient and Accurate Video Large Language Models via Context Mining
Abstract
How can we minimize memory consumption during Video Large Language Model (VLLM) inference while preserving temporal reasoning accuracy? VLLMs achieve strong temporal reasoning performance in video understanding by processing rich visual information across video frames, yet retaining the full visual representation incurs considerable memory overhead. To alleviate this overhead, existing approaches compress video representations by selectively retaining only visual information deemed important. Specifically, they excavate information associated with key events. However, these approaches do not consider preserving the visual evidence that temporally precedes such events, even though this evidence provides critical context for interpreting temporally subsequent key events. Consequently, aggressive compression degrades temporal reasoning accuracy in VLLMs. In this paper, we propose CueMiner, an efficient video compression method that preserves temporal reasoning accuracy even under aggressive compression. To preserve accuracy, CueMiner not only captures salient events, but also mines the temporally antecedent visual context that enables VLLMs to interpret how subsequent events develop over time. On MLVU with LLaVA-OneVision, CueMiner retains 99.4% of the accuracy at a retention ratio of 5%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.