Refine Before You Regress: Persistent Belief Refinement for Geometric Reasoning with Frozen VLMs
Abstract
Iterative geometric reasoning raises a basic design question: should the numerical answer also serve as the state for subsequent refinement? We study this question through full-scene projective cuboid prediction from grayscale observations of a single RGB image with frozen vision–language models (VLMs). We introduce persistent belief refinement, a training-free inference protocol that maintains persistent object identities and revisable hypotheses about object support, shape, camera and scene relations, and uncertainty, while generating cuboid coordinates only at the final readout. A recurrent-sufficiency analysis motivates this separation by characterizing when a terminal output can replace a richer state without changing subsequent task-level behavior. On 90 held-out SUN RGB-D scenes, the protocol improves coverage-aware Cuboid-F1 AUC by under Raw observations and under coarse-to-fine observations relative to coordinate regeneration, with both protocols using six logical calls. Both configurations also substantially reduce false positives. On a 20-scene subset, adding coordinate feedback while retaining the belief-state schema and persistence rules yields lower F1 at unchanged aggregate recall under approximately matched token usage. A 17-scene cross-backbone comparison finds higher Cuboid-F1 AUC for the complete belief-refinement protocol in all three tested VLM families. These findings support separating recurrent hypothesis refinement from terminal numerical prediction in frozen-VLM cuboid inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.