acceptodds
Under review as a conference paper at ICLR 2027

CueMiner: Efficient and Accurate Video Large Language Models via Context Mining

Abstract

How can we minimize memory consumption during Video Large Language Model (VLLM) inference while preserving temporal reasoning accuracy? VLLMs achieve strong temporal reasoning performance in video understanding by processing rich visual information across video frames, yet retaining the full visual representation incurs considerable memory overhead. To alleviate this overhead, existing approaches compress video representations by selectively retaining only visual information deemed important. Specifically, they excavate information associated with key events. However, these approaches do not consider preserving the visual evidence that temporally precedes such events, even though this evidence provides critical context for interpreting temporally subsequent key events. Consequently, aggressive compression degrades temporal reasoning accuracy in VLLMs. In this paper, we propose CueMiner, an efficient video compression method that preserves temporal reasoning accuracy even under aggressive compression. To preserve accuracy, CueMiner not only captures salient events, but also mines the temporally antecedent visual context that enables VLLMs to interpret how subsequent events develop over time. On MLVU with LLaVA-OneVision, CueMiner retains 99.4% of the accuracy at a retention ratio of 5%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.