acceptodds
Under review as a conference paper at ICLR 2027

SRET: Spatial Relation-Aware Evidence Transport for Visual Token Reduction

Abstract

Visual token reduction is key to lowering vision language inference costs, but redundancy in object semantics does not imply redundancy in relational evidence. Spatial reasoning relies on information distributed across objects, reference landmarks, and observations. Importance-based deletion or semantic merging may preserve object appearance while disrupting the structure needed to infer direction, position, and viewpoint changes. We propose SRET, a training-free method for spatial relation-aware evidence transport. Inspired by geometric constraints on spatial observations, SRET formulates compression as relational representative selection and evidence transport. Visual semantics, shallow context, and local layout define token relations, guiding coverage repair and cross-group budget adjustment under a fixed budget. Based on source–anchor relations, source evidence mass is split into a node component and a relational residual, both transported to structure-aware receivers. Selection and aggregation share evidence weights, while tangential updates preserve anchor norms. Experiments on two vision language models and multiple spatial reasoning benchmarks demonstrate that SRET preserves spatial reasoning ability while compressing visual sequences.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.