Binding What to Where: Configurational Binding via On-Policy Self-Distillation for Multi-Object Spatial Reasoning
Abstract
Vision-Language Models (VLMs) excel at visual understanding, but still struggle to reason reliably about spatial relationships. Models may capture relevant object cues while failing to bind them into a globally consistent spatial configuration, limiting their ability to reason about relations across multiple objects. Models may capture relevant visual cues while failing to organize them into a coherent spatial representation, limiting their ability to reason reliably about spatial structure. To address this, we propose BindOPSD, a post-training framework that strengthens spatial reasoning through configurational binding. BindOPSD first constructs a privileged teacher context with a joint spatial view and correspondence-aware attention, which organizes local object evidence within a shared spatial frame. It then transfers this capability to the student through on-policy self-distillation, enabling the model to internalize the enhanced spatial reasoning while requiring only the original image at inference. To preserve general visual capabilities during spatial adaptation, we further introduce gradient decomposition to reduce interference between spatial learning and general capability retention. We also introduce FineBind, a fine-grained spatial reasoning benchmark designed to evaluate global configuration and object-to-position binding. Experiments across diverse spatial reasoning benchmarks demonstrate consistent improvements over strong baselines, with effective transfer beyond the training distribution and no additional visual inputs at inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.