Same Object, or Its Twin? Diagnosing Instance Association Across Scene Revisits
Abstract
Persistent object maps must maintain physical identity across changes in position and viewpoint. An incorrect association can reflect absent observational cues, difficulty reading them from features, or limitations in how matching uses them. We test these explanations through simulated paired revisits of otherwise identical chairs with artificial persistent color markers. Representations share oracle masks, views and the full candidate map. In the primary strong-color swap condition with correlated cameras, an RGB descriptor recovers most identities. A development-fitted linear probe decodes marker colors from MASA-R50 features, yet maximum view-pair cosine matching frequently selects the exchanged partner. Holding view features and candidates fixed, normalized-mean aggregation increases correct associations from 2 to 10 of 36 queries. Penalizing differences in the probe's continuous output raises this count to 15. Aggregation accounts for most of both this recovery and the accuracy loss on other map objects, while readout gains vary across representations and camera conditions. This partial recovery shows that evaluating readable identity cues requires testing how fixed features are aggregated and compared under full-map candidate competition.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.