CLAP: Classification and Localization with Object-Aligned Prototype Evidence
Abstract
Part-prototype networks make image classification transparent by expressing predictions through spatial matches to learned prototypes. Yet classification supervision rewards any discriminative cue, so prototypes may respond to a small object part or correlated background rather than faithfully represent the target object. This limits the use of their spatial responses for object localization. We introduce CLAP, a training framework that makes prototype evidence object-aligned using only image-level labels. A lightweight class-agnostic foreground–background head produces pseudo-masks, and two alignment losses encourage collective coverage of the predicted foreground while suppressing responses on predicted background. Classification and localization are then computed from the same class-associated spatial prototype responses, allowing the prototype pathway to retain its discriminative role while gaining localization capability. We instantiate CLAP with ProtoPNet and ProtoPool. On part-prototype benchmarks, CLAP substantially improves GT-Known localization while preserving classification accuracy with no additional cost at inference. On standard WSOL benchmarks, it achieves 97.5% GT-Known localization on CUB-200-2011 and 74.9% on ILSVRC, remaining competitive with specialized WSOL methods. These results show that part prototypes can provide a unified, intrinsically traceable basis for recognition and localization, establishing prototype-based WSOL as a promising direction for future part-prototype architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.