AdaDPC: Adaptive Dual-Path Computation for Efficient Long-Video Understanding
Abstract
Long-video understanding requires models to process large visual token sequences, yet useful visual evidence does not necessarily require full updates at every layer. The key challenge is therefore to reduce redundant visual state refinement without losing access to information that may still support the answer. We propose AdaDPC, an adaptive dual-path computation framework that explicitly separates visual state refinement from visual evidence access. The Selective Refinement Path performs full attention and FFN updates on dynamically selected visual blocks, while the Query-Guided Access Path allows text queries to retrieve evidence from the remaining visual states at low cost. Both paths operate on a shared evolving visual state sequence, allowing visual blocks to change computational roles across layers. Cross-Layer Priority Feedback further uses question-relevant evidence discovered during access to guide subsequent refinement within a fixed budget. Across eight benchmarks, AdaDPC outperforms the dense model, improving MLVU Dev from 65.3 to 69.7 and TimeScope from 76.3 to 81.9, gains of 4.4 and 5.6 percentage points, respectively. Meanwhile, AdaDPC reduces FLOPs by 55.5%, demonstrating that broad visual evidence access can be preserved without uniformly applying expensive refinement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.