acceptodds
Under review as a conference paper at ICLR 2027

When the Agent Turns: Correctness-Aware Multimodal Reasoning over Panoramic Yaw Orbits

Abstract

When an embodied agent turns, the scene remains unchanged, but the correct answer to an egocentric spatial question may change. We identify a failure in panoramic vision–language models (VLMs): stable view-averaged accuracy can coexist withstale-frame predictions that retain the source answer when the reference frame requires an update. To measure this failure, we introduce a correctness-aware yaw-orbit benchmark that preserves image content through cyclic shifts and distinguishes invariant questions from equivariant questions whose targets are regenerated from 3D geometry. Across five open VLMs, mean transformed accuracy on invariant questions remains within 3.25 points of clean accuracy, yet full-orbit accuracy is 24.79–35.12% and changed-label accuracy is 16.25–30.58%. We propose OrbitAlign, a post-training framework that incorporates the answer transformation law into supervision and optimization. Orbit Supervised Fine Tuning (SFT) learns from geometry-regenerated targets; orbit-based Group Relative Policy Optimization (GRPO) jointly optimizes paired headings; and a correctness-gated Frame-Canonical Orbit (FCO) reward couples predictions only when both are correct in their respective frames. Across three backbones, OrbitAlign improves held-out changed-label accuracy over PanoEnv GRPO by 32.05–42.17 percentage points. With 100 Reinforcement Learning (RL) steps following orbit SFT, the complete framework reduces H200 training GPU-hours by a factor of 4.1–5.2. Ablations examine the contributions of target regeneration, orbit grouping, and cross-view correctness coupling. Together, these results establish current-frame correctness as a criterion for both evaluating and training panoramic spatial reasoning models. The benchmark, code, and model checkpoints will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.