acceptodds
Under review as a conference paper at ICLR 2027

PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization

Abstract

3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB–point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ with how the object appears in the search frames. The benchmark baseline predicts cuboids after feature modeling, leaving their geometry unused for subsequent feature refinement. We investigate whether complete intermediate cuboids can improve query and proposal representations before final decoding. To this end, we introduce the Progressive Geometric Learning for 3DVQL (PGL-3D), a predict–select–refine framework that uses complete intermediate cuboids to guide the aggregation of search evidence and update both query and proposal representations. A shared head first predicts a complete cuboid for every proposal. Query–Tube–Memory (QTM) then selects reference observations by combining proposal association, cuboid quality, frame response, and target absence, because association confidence alone establishes neither target presence nor geometric accuracy. The center, size, and orientation of each selected cuboid define soft pooling weights over the query-conditioned proposal features of its frame. The pooled memory updates both the query and the proposal representations, and the head re-predicts from the updated features. A training-only objective, ST-D9O, supervises the complete cuboid geometry at every stage by adding boundary, signed-distance, and soft-overlap terms to parameter regression. PGL-3D achieves a mean stAP of 0.2687 ± 0.0062 on 3DVQL, compared with 0.044 reported for LaF. Component ablations support the benefits of geometry-guided feature updates beyond intermediate supervision, while stage-wise analyses show improved cuboid accuracy. Replacing the geometry objective in our PROT3D reproduction with ST-D9O further improves mAO on GSOT3D from 21.63% to 25.78%. Our code and models will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.