acceptodds
Under review as a conference paper at ICLR 2027

READ: Reading the Selective Scan for Faithful Video Attribution

Abstract

Selective state-space models are an efficient approach to long-sequence video modeling. However, their recurrent computation makes it difficult to determine which video tokens drive a prediction. Existing attribution methods often require additional backward passes or repeated model evaluations, making them substantially more expensive than the forward inference they aim to explain. In this work, we introduce READ, a parameter-free attribution method for selective state-space models that unrolls the recurrence at the position read by the classifier. READ traces the contribution accumulated in the recurrent state back to the video tokens that contribute to the final prediction. By anchoring attribution at the classifier's readout, it produces a single spatiotemporal relevance map from which frame- and region-level relevance can be obtained directly. Despite requiring only one forward pass, READ achieves competitive faithfulness and localisation compared with substantially more expensive attribution methods. Experimental results show that READ avoids the systematic relevance biases of existing methods while recovering meaningful temporal and class-specific structure across datasets and backbones. More broadly, our results show that the recurrence underlying selective state-space models is not only an efficient mechanism for inference, but also a direct source of interpretable evidence for their predictions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.