GLANCE: Free-Form 3D Object Identification through Active Geometric Inspection
Abstract
Identifying an arbitrary 3D object from geometry alone is challenging: a single viewpoint may be insufficient, and geometric identity need not admit a unique textual ground truth. We study this problem as free-form 3D object identification: inferring an object’s identity without color, texture, physical scale, canonical orientation, scene context, or a candidate vocabulary. We introduce GLANCE, an active recognition approach in which an off-the-shelf multimodal large language model (MLLM) iteratively inspects rendered geometric observations. Rather than committing to an initial guess, GLANCE maintains competing free-form identity hypotheses, predicts geometric evidence that would distinguish them, and selects subsequent viewpoints to test these predictions. No task-specific 3D training or candidate vocabulary is required. To evaluate free-form identification without imposing a predefined taxonomy, we construct a benchmark of 1,118 3D shapes, each independently identified by 4 human annotators together with their confidence score. We retain these annotations as multiple reference identities rather than collapsing them into a single ground truth, and evaluate predictions through semantic agreement with the references. In controlled experiments, GLANCE improves recognition over passive multi-view baselines while requiring fewer observations. At larger scale, GLANCE achieves 83.2% primary accuracy on 570 benchmark objects. On ShapeNet, GLANCE achieves 72.0% accuracy while keeping the category vocabulary hidden from the recognition model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.