MaxMoreGaze: Max-Loss and More-Salient Gazing for Efficient Video Understanding
Abstract
Efficient video understanding must preserve spatiotemporal features within a limited visual-token budget. Reducing the frame rate discards temporal structure, and lowering resolution removes small objects, text, and other local details. Reconstruction-guided selectors acquire multi-scale patches before full vision encoding, yet a frame-average reconstruction target conceals local failures and yields unstable stopping. We present **MaxMoreGaze**. **MaxGaze** replaces that scalar with a bias-corrected 7 × 7 reconstruction-error map: the same map ranks the next patch, constructs the training trajectory, and decides when to stop, namely when every spatial window meets the target. **MoreGaze** then uses gray-calibrated cross-scale attention from the frozen vision encoder to prune weak fine-scale patches in a training-free manner. Retained patches are serialized with explicit tile and frame structure. At matched patch budgets, MaxGaze improves local reconstruction and stopping. On seven video-understanding settings with identical frames and tiles, MaxMoreGaze improves the performance at a lower token cost than the previous method, and is stronger on detail-intensive high-resolution benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.