acceptodds
Under review as a conference paper at ICLR 2027

SGT-3D: 3D-Guided Self-Distillation and Geometric Verification for Spatial Reasoning

Abstract

Spatial reasoning from images often involves estimating object geometry and using it to answer spatial questions. Some spatial reasoning approaches expose intermediate 3D attributes, but making them visible does not ensure accurate prediction or correct use. These two goals motivate complementary supervision: final-answer rewards do not directly assess attribute accuracy, while training on reference responses may leave reasoning from the model's own imperfect predictions insufficiently supervised. The resulting challenge is to improve object-level spatial understanding and turn better estimates into more accurate answers. We present , which couples learning to predict object centers, dimensions and orientations with learning to reason from the states the student actually generates. A training-only 3D expert guides self-distillation, with continuation supervision following the student's stated attributes within a fixed-state interface, and subsequent reinforcement learning supplies numerical attribute feedback alongside answer rewards. Inference requires only RGB images and questions. SGT-3D achieves QA accuracy of 69.2% on 3DSRBench-real and 96.2% on CV-Bench-3D, with 75.0% simultaneous center/dimension/orientation pass rate on the separate SUN-G400 audit using native SUN RGB-D annotations and fixed object-region prompts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.