acceptodds
Under review as a conference paper at ICLR 2027

PanoAgent: Efficiently Understanding in 360° through Active Inspection

Abstract

Pretrained vision-language models (VLMs) exhibit strong perception and reasoning capabilities on perspective images, but panoramic projection distortions limit the use of these capabilities in panoramic understanding. Existing adaptation methods primarily rely on fine-tuning with panoramic question–answer supervision or predefined workflows that convert panoramas into planar views for the model. We introduce PanoAgent, which adapts VLMs to panoramic understanding by learning how to observe. It jointly learns to select question-relevant panoramic transformations and generate answers, allowing the model to actively select observations suited to its existing visual capabilities. Specifically, the model uses panorama-specific tools to select local planar views or update the global reference heading as needed, with returned evidence guiding subsequent observation and reasoning. Trajectory demonstrations connect panoramic concepts to observation actions, while reinforcement learning further refines the policy through outcome and geometric feedback. Experiments across multiple panoramic benchmarks show that PanoAgent achieves leading overall performance and offers a better balance among understanding accuracy, data efficiency, and inference efficiency than existing panoramic adaptation methods. The dataset and trained models will be released publicly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.