acceptodds
Under review as a conference paper at ICLR 2027

Whose View Wins? Inside Egocentric Bias in Visuospatial Perspective Taking

Abstract

Visuospatial perspective taking (VSPT) requires reasoning from a viewpoint different from the viewer’s own, yet Multimodal Large Language Models (MLLMs) often exhibit egocentric bias. We investigate the representational mechanism underlying this failure using controlled scenes that independently manipulate egocentric and altercentric spatial relations. Linear probing shows that both relations are decodable from hidden representations despite poor VSPT performance under perspective conflict, suggesting that altercentric information is not simply missing. We then decompose hidden representations into egocentric and altercentric components and quantify their contributions along the ego–alter decision direction. Under conflict, the egocentric component contributes systematically more strongly toward its corresponding answer than the altercentric component. This imbalance is accompanied by a large and consistent egocentric advantage in component magnitude, while alignment differences are weaker and model-dependent. Guided by this analysis, we intervene on magnitude and alignment separately; both improve conflict-case performance, with magnitude suppression generally yielding larger gains. The decomposition also transfers to natural-image benchmarks in several settings. Together, our results suggest that egocentric bias stems from an egocentric-favoring imbalance in how concurrently represented perspectives contribute to model predictions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.