Evi-Prefill: Evidence-Aware Prefill for Efficient Long Video Understanding
Abstract
Existing acceleration methods for long video large models adopt a paradigm that uniformly processes video tokens. This ignores the sequential arrival of task-relevant : once sufficient evidence emerges early, models still process the full video stream, leading to redundant computation. To extend video temporal coverage, chunk-wise inference adopts partial observability yet remains . In this paper, we present , an framework allocating subsequent compute from visual evidence accumulated under partial observability. Specifically, we formulate evidence gathering for incoming video chunks as a Partially Observable Markov Decision Process. With this model, we estimate evidence demand via the model’s internal signals and actively switch from inference, reducing redundant compute while preserving performance. Moreover, Evi-Prefill filters structurally redundant frames before visual encoding to boost evidence estimation efficiency. Evi-Prefill requires no training and complements prior methods. Evaluations over 3 benchmarks achieve 1.5–1.9 speedup without average performance degradation, opening a new route to efficient long video understanding via active visual-evidence acquisition. Code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.