Location: What View Manifolds Can and Cannot Tell Us About Invariant Recognition
Abstract
Deep vision networks could recognise certain objects from every viewpoint well but fail on others, and we hypothesise two possible ways in which recognition invariance is achieved through the viewpoint manifold of the object: gradual collapse in later layers, or reorganisation. We test both possibilities against the networks' classification heads and found three factors influencing readout consistency, which are how far the view set sits from the decision boundary, how widely they spread, and how much of the spread the readout measures on the viewpoint manifold. We measured the factors using 15 ShapeNet objects rendered at forty angles in a full azimuth rotation with ViT-B/16 and ResNet-50. In the two networks tested, correct recognition is rarely achieved, with two and five in fifteen objects, respectively. Relative collapse of the manifold across depth is not a requirement for correct recognition either. Compactness of viewpoints is positively related to the consistency of the recognition and shows no detectable association with its correctness. Correctness is related to where the viewpoint set lies relative to the boundary. Summaries computed only from distances between views cannot certify recognition.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.