Spatial Bias in Vision Language Models
Abstract
Spatial intelligence is essential for embodied tasks such as navigation, manipulation, and agentic interaction. Vision-Language Models (VLMs) increasingly exhibit spatial reasoning capabilities, yet priors learned from large-scale vision-language pre-training may override current visual evidence and induce systematic spatial errors. While social bias and visual hallucination have been widely studied, spatial bias remains underexplored. We define spatial bias as the systematic tendency of VLMs toward prior-aligned spatial predictions despite conflicting visual evidence. To investigate this phenomenon, we construct an original–counterfactual embodied visual question answering (VQA) benchmark covering diverse spatial tasks and prior sources, and introduce a multi-stage diagnostic framework that distinguishes counterfactual failures from prior-retention bias, language-aligned bias, and non-bias errors. We further investigate how input organization, contrastive context, background information, and visual guidance affect spatial bias, and evaluate multiple mitigation strategies. Experiments on diverse VLMs reveal widespread spatial bias that varies across model family, spatial task, prior source, and contextual conditions such as memory and visual information. Input and background configurations affect counterfactual robustness, while selective context suppression and explicit visual guidance can reduce bias. Prompt-based interventions yield limited and inconsistent gains, whereas supervised fine-tuning provides more reliable mitigation. These findings underscore the limitations of accuracy-only evaluation and the need to explicitly diagnose and mitigate spatial bias in embodied VLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.