Transport What Matters: Evidence-Directed Optimal Transport for Long Video Understanding
Abstract
Long video understanding often requires reducing thousands of observations to a small frame budget that a vision–language model can process. Existing selectors prioritize query relevance, diversity, or adaptive sampling, but leave implicit how evidence from discarded observations should be represented by the retained set. We introduce Evidence-Directed Optimal Transport (E-DOT), a training-free framework that casts frame selection as transporting task-conditioned evidence mass from dense video observations to a sparse support of frames. Query relevance over dense observations and temporal extent define the source mass. A directed ground cost combines temporal proximity with sparse local-helpfulness estimates from a frozen model: transport to a nearby frame that preserves or improves source evidence is cheap, whereas losses in relevance or helpfulness are penalized. Relaxing the destination marginal lets multiple sources share representatives, allowing the support to concentrate on evidence that remains poorly covered. The resulting support-selection objective admits monotone-submodular greedy optimization and coverage characterizations. Across four long-video question answering benchmarks and three answerers, E-DOT outperforms all compared frame selectors under a fixed answerer budget, while matched ablations isolate the contributions of local helpfulness, directed transport geometry, and coverage-aware support selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.