When Retrieval Attribution Misses Visual Evidence in Late-Interaction Models
Abstract
Does higher annotation overlap in a retrieval-score highlight reflect sensitivity to the matched query? We test both properties in late-interaction visual retrievers. At , and , fixed query-bank correction increases mean max-direct RegionHit@4 in all ten original checkpoint–support conditions while retaining retrieval scores. An industrial-source position prior nevertheless has higher full-support means than corrected max-direct in all ten, with seven contrasts surviving its separate Holm10 procedure, conditional on the fitted field. A retrospective same-page analysis on a separate subset of 503 eligible pages compares archived highlights against the annotations of their own and another real question. Corrected max-direct favors matched annotations in all ten means, by 2.31 to 19.13 percentage points; seven contrasts meet the adjusted threshold in the new Holm20 family. Correction increases this advantage in all ten means, with six increments meeting that threshold. Inference on this eligible annotation subset is limited by few, unequal dependency blocks. A supporting v0.2 audit traces the suppression–deletion gap to contribution transfer. The findings show why localization levels and matched-query annotation advantages should be reported separately; answer utility remains unestablished.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.