acceptodds
Under review as a conference paper at ICLR 2027

MaRR: Map-Based Rotation Reasoning for Vision–Language Models

Abstract

Forward mental rotation requires predicting an object’s unseen pose after a specified rotation, rather than recognizing a visible configuration. Our controlled diagnostics show that frozen vision– language model features contain decodable rotation information and support a supplied transformation, while inferring and reusing transformations across objects remains unreliable. Building on this insight, we introduce MaRR (Map-based Rotation Reasoning), a teacher-guided self-distillation framework that combines object- adaptive representations with a shared, object-independent transformation procedure. Using teacher traces, MaRR teaches VLMs to map object-specific features, rotate their spatial relations, and match the predicted configuration to candidate images. Fine- tuning Qwen3.5-4B with this supervision improves held-out CAD accuracy, with gains extending to shifted CAD renderings and block objects, although transfer to public benchmarks remains selective. A matched comparison shows that trace design introduces tradeoffs: natural teacher responses yield stronger OOD transfer and physical-choice consistency, while standardized MaRR traces achieve higher accuracy under candidate permutation. Together, these findings suggest that structured supervision can help VLMs organize and use their latent spatial knowledge, while the way the procedure is expressed in teacher traces affects transfer and robustness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.