acceptodds
Under review as a conference paper at ICLR 2027

Models Already Know Where to Look: Recovering the Discarded Spatial Map in Parallel Scanpath Prediction

Abstract

Task-driven scanpath prediction models an ordered, variable-length sequence of fixations made during a visual task. Some parallel models supervise a spatial distribution for each fixation slot but discard it at inference, using an independent coordinate-regression head instead. We study this representation–decoding mismatch through a controlled family of models. C0 uses coordinate regression, S0 adds an auxiliary spatial map, and S0-SpatialDecode (S0-SD) decodes the map of an already trained S0 checkpoint without retraining. Task-Conditioned Spatial Decoding (TCSD) instead trains and predicts through the same spatial distribution. On an image-disjoint COCO-Search18 evaluation, S0-SD improves all six metrics over S0 under joint seed–image resampling. A stronger regression head narrows but does not close the gap to spatial decoding. S0-SD and TCSD show no reliable difference on any metric, indicating that most of the gain comes from using the learned map at inference rather than from the TCSD training pathway. The effect persists with ImageNet-pretrained features, while experiments on Gazeformer and AiR-D show that its transfer is metric- and task-dependent. TCSD retains efficient one-pass inference, although HAT-SF remains more accurate on most metrics. These results show that spatial supervision and the decoder that turns the learned map into fixation coordinates should be evaluated together.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.