acceptodds
Under review as a conference paper at ICLR 2027

GAUGE3D: COORDINATE-GAUGE CONSISTENCY FOR DENSE 3D VISION–LANGUAGE GROUNDING

Abstract

Geometry-aware vision–language models ground language in 3D scenes, yet high accuracy in one coordinate frame can conceal sensitivity to arbitrary coordinate choices. Changing the global origin or orientation preserves the observed scene but requires dense masks, grounded text spans, and language answers to follow consistent transformation rules. We introduce , a framework for evaluating and learning coordinate consistency in dense 3D vision–language grounding. The key idea is to turn coordinate re-expression into paired supervision of the complete output interface. A geometry-only compiler transforms camera poses, world-space coordinates, and geometric targets while preserving the underlying color and depth observations. A joint objective constrains mask equivariance on a common source support, text-span grounding consistency, and answer invariance for frame-independent queries, with transformed targets for frame-dependent queries. A consistency audit activates paired training when repeatable discrepancies exceed the identity and compiler noise floor, linking diagnosis to targeted supervision without adding a new 3D encoder. Evaluations on ScanRefer, ScanQA, and SQA3D demonstrate that improves both dense grounding and spatial question answering over the underlying 3D vision–language model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.