PanoAgent: Efficiently Understanding in 360° through Active Inspection
Abstract
Pretrained vision-language models (VLMs) exhibit strong perception and reasoning capabilities on perspective images, but panoramic projection distortions limit the use of these capabilities in panoramic understanding. Existing adaptation methods primarily rely on fine-tuning with panoramic question–answer supervision or predefined workflows that convert panoramas into planar views for the model. We introduce PanoAgent, which adapts VLMs to panoramic understanding by learning how to observe. It jointly learns to select question-relevant panoramic transformations and generate answers, allowing the model to actively select observations suited to its existing visual capabilities. Specifically, the model uses panorama-specific tools to select local planar views or update the global reference heading as needed, with returned evidence guiding subsequent observation and reasoning. Trajectory demonstrations connect panoramic concepts to observation actions, while reinforcement learning further refines the policy through outcome and geometric feedback. Experiments across multiple panoramic benchmarks show that PanoAgent achieves leading overall performance and offers a better balance among understanding accuracy, data efficiency, and inference efficiency than existing panoramic adaptation methods. The dataset and trained models will be released publicly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.