EgoAudit: Auditing Egocentric Spatial Reasoning in 3D Vision-Language Models
Abstract
3D vision-language models (3D VLMs) have made rapid progress in grounding objects and answering questions in 3D indoor scenes. Many queries that people pose are egocentric: "Which lamp is on my left?" has a definite answer only once the model knows where the speaker stands and which way they face. Existing benchmarks and models often provide this viewpoint as an annotated pose, so a high score may reflect how well a model uses a given viewpoint rather than whether it can infer the viewpoint from the description. We present EgoAudit, a framework for auditing egocentric spatial reasoning in 3D VLMs. EgoAudit equips a 3D VLM with a viewpoint-conditioned object encoding, through which the viewpoint can be removed, perturbed, or replaced while the scene, the query, and the model stay fixed, and it evaluates the model with three protocols: viewpoint intervention, matched training with and without the viewpoint, and substitution of the viewpoint source. Extensive experiments on public 3D grounding and 3D visual question answering benchmarks show that the audited model, a strong 3D VLM, follows a supplied viewpoint closely but does not infer it: much of its advantage on egocentric queries is lost when the annotated viewpoint is withheld or replaced by the estimate of a learned localizer. These results indicate that scores obtained with annotated viewpoints measure viewpoint use rather than perspective taking, and we propose a reporting protocol for egocentric 3D evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.