ARK: A Benchmark for Reasoning Re-Identification in Multimodal Large Language Models
Abstract
Re-identification (ReID) is commonly formulated as a retrieval problem, yet ranking alone does not always resolve identity: an observation of the correct identity may appear among the top-k candidates without being ranked first, while open-set queries require recognizing when no valid match exists. We formulate this unresolved post-retrieval stage as Reasoning ReID, where a multimodal large language model (MLLM) verifies retrieved candidates before producing the final identity decision. To systematically study the capabilities required for such verification, we introduce ARK, a benchmark for Reasoning ReID using Animal ReID as a challenging testbed. ARK spans 17 species and seven controlled protocols organized into three capability families: Visual Identity Perception, Evidence Integration & Association, and Decision Reliability. An evaluation of 21 proprietary and open-weight MLLMs alongside a human reference shows that current models can benefit from complementary evidence, while their reliability remains uneven under demanding verification conditions. We further instantiate Reasoning ReID in retrieval pipelines built on CLIP and SigLIP, showing that selectively invoking MLLM verification for uncertain retrievals can improve end-to-end identification in both closed- and open-set settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.