PanoArena: Zero-Shot Benchmark for Panoramic Understanding
Abstract
Existing evaluations of panoramic understanding remain largely image-centric and fail to systematically assess panorama-specific properties, particularly view-dependent spatial reasoning under the complete 360 field of view (FoV). To bridge this gap, we introduce PanoArena, a comprehensive benchmark with 4,631 samples for evaluating general-purpose panoramic understanding across both images and videos. PanoArena adopts a unified question-answering format with more than 10 sub-tasks spanning semantic and spatial understanding, where questions are further categorized into bearing-dependent and bearing-invariant settings based on whether the observer's view direction is required for reasoning. We systematically benchmark existing multimodal large language models (MLLMs) across different panoramic representations under zero-shot settings. We further propose PanoCompass, a training-free method, enabling a cleaner assessment of intrinsic panoramic reasoning rather than task-specific memorization. Extensive experiments reveal substantial limitations of current MLLMs, particularly in bearing-dependent spatial reasoning, providing new insights and directions for general-purpose panoramic understanding. The benchmark data and relevant code will be made available to the public.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.