Content-Aware Scanning in Vision State Space Models: Is Grouping Enough?
Abstract
Content-aware scanning for vision state space models (SSMs) rests on a proximity hypothesis: placing informative tokens next to each other shields their interaction from the recurrence's decay. We test it in a controlled SSM restoration testbed (1,108 trained models) that varies the grouping of informative tokens independently of the information an order carries, and the hypothesis fails in both directions. Sorting tokens by training-free importance scores, such as constant false alarm rate (CFAR) detection, raises grouping without adding information and performs like a random permutation (within 0.1 dB), even when the scores find the targets with AUC 0.98. Ground-truth orders that place informative tokens farther apart than a raster scan still help, in all 72 runs, so "oracle" orders measure leakage rather than an attainable gain. In a factorial design that sets the precision of an order (the informative fraction of the tokens it selects) and its grouping independently, raising precision from 0 to 1 adds 3.5 dB at fixed grouping, and grouping helps only in proportion to precision: grouping alone explains 0.62 of the variance, grouping and precision with their interaction 0.96. The same information delivered as an input channel instead of an order helps as much or more (+9.82 against +8.03 dB for the ground-truth support; +0.78 against +0.69 dB for a supervised CNN scorer). Finally, because the recurrence has no history at the start of a scan, position is readable there: a ground-truth block placed first gains 3–6 dB more than mid-sequence in a causal scan, also with the scan convolution removed and with a variable number of targets. We release a protocol with region-aware metrics and random-permutation, locality, leakage, and input-channel controls.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.