MomentRoute: Asymmetric Resolution Routing for Budgeted Visual Memory
Abstract
Budgeted visual memory must search a broad video history while preserving the visual detail needed to answer a query. We show that these two tasks have different resolution requirements: across two cross-modal rerankers, reducing candidate resolution from to lowers selection recall by only 0.23–0.26 percentage points, whereas reducing three selected frames to lowers the 72B reader’s conditional answer accuracy by 10.4 points. MomentRoute exploits this asymmetry by separating candidate localization from evidence reading: it reranks 32 thumbnails in the cloud, then requests three frames for answering. On 1,200 constructed Ego4D-NLQ QA episodes, this allocation reaches 65.1% accuracy, compared with 60.1% for direct thumbnail upload and 58.4% for compressed full-buffer video. Relative to adaptive edge selection with three high-resolution frames, cross-modal reranking improves QA by +5.2 pp (). Accuracy–resource curves expose the accompanying trade-offs: four-frame edge selection reaches 64.1% at 1,009 KiB, compared with 65.1% at 1,136 KiB for two-stage routing (exact McNemar ). Edge selection provides lower TTFT; two-stage routing uses 25% fewer reader visual tokens (768 vs. 1,024). Together, these measurements show how allocating resolution by processing stage preserves broad candidate search and detailed evidence reading, while making the communication and inference trade-offs explicit. Code and benchmark data will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.