acceptodds
Under review as a conference paper at ICLR 2027

MomentRoute: Asymmetric Resolution Routing for Budgeted Visual Memory

Abstract

Budgeted visual memory must search a broad video history while preserving the visual detail needed to answer a query. We show that these two tasks have different resolution requirements: across two cross-modal rerankers, reducing candidate resolution from to lowers selection recall by only 0.23–0.26 percentage points, whereas reducing three selected frames to lowers the 72B reader’s conditional answer accuracy by 10.4 points. MomentRoute exploits this asymmetry by separating candidate localization from evidence reading: it reranks 32 thumbnails in the cloud, then requests three frames for answering. On 1,200 constructed Ego4D-NLQ QA episodes, this allocation reaches 65.1% accuracy, compared with 60.1% for direct thumbnail upload and 58.4% for compressed full-buffer video. Relative to adaptive edge selection with three high-resolution frames, cross-modal reranking improves QA by +5.2 pp (). Accuracy–resource curves expose the accompanying trade-offs: four-frame edge selection reaches 64.1% at 1,009 KiB, compared with 65.1% at 1,136 KiB for two-stage routing (exact McNemar ). Edge selection provides lower TTFT; two-stage routing uses 25% fewer reader visual tokens (768 vs. 1,024). Together, these measurements show how allocating resolution by processing stage preserves broad candidate search and detailed evidence reading, while making the communication and inference trade-offs explicit. Code and benchmark data will be made publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.