acceptodds
Under review as a conference paper at ICLR 2027

SoccerCanon: A challenging Benchmark and the Case for Stitching Vision Experts over VLMs

Abstract

Soccer player recognition in broadcast imagery remains challenging under blur, occlusion, small scale, and extreme viewpoints, where identity cues such as faces, appearance, and jersey numbers are often incomplete or unreliable. Existing benchmarks provide limited support for systematically evaluating recognition under such heterogeneous visual evidence. To address this limitation, we introduce SoccerCanon, a benchmark of 1,108 labeled player instances, comprising 613 validation and 495 test instances. Face recognizers that identify clean player photographs with about 96% R@1 reach only 40-47% on SoccerCanon, and accuracy collapses on the small faces that dominate broadcast imagery. We evaluate specialized vision models and general-purpose vision-language models (VLMs) under controlled, modality-aware settings. Our experiments reveal that different visual cues exhibit complementary strengths, while no single recognition pathway remains consistently reliable across challenging cases. Motivated by this observation, we propose a training-free expert-stitching framework that combines face, appearance, and number experts according to their query-specific reliability. Our framework reaches 60.94% R@1 on the face-available test observations without match metadata, whereas the strongest VLM reaches 29.24% even when given match-restricted candidates, and reliability-aware stitching improves most where face evidence is weak.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.