acceptodds
Under review as a conference paper at ICLR 2027

LR-Bench: Diagnosing and Mitigating Left-Right Confusion in MLLMs for Human-Centric Spatial Understanding

Abstract

Human subject-centered left-right (LR) discrimination, i.e., answering “which is his left hand” in the depicted person's own frame of reference (FoR) rather than in screen coordinates, is a fundamental capability for human-centric vision-language applications. We find that current multimodal large language models (MLLMs) fail at this task in a systematic way. Most open-source models score near chance, and for most mirror pairs their prediction stays the same after the input is horizontally flipped, which means the errors are systematic rather than random. To diagnose this failure, we introduce LR-Bench, a benchmark of 8,048 human-annotated binary QA pairs over 2,942 in-the-wild images and 1,809 video clips, spanning appearance, action & pose, spatial relation, and temporal dynamics, and evaluated with binary accuracy, a mirror-paired consistency score, and an LLM-judged caption protocol. Beyond evaluation, we propose Spatial Latent Prompt (SLP), a plug-in that injects an input-derived human structural prior, a Mesh-Derived Part Map from an off-the-shelf mesh recovery model, into frozen visual tokens through learnable part embeddings and a person-relative positional channel. With fewer than 0.1M trainable parameters and a frozen backbone, SLP improves three MLLM backbones by 36.9–48.7 points on images and 27.9–35.8 points on videos, and clearly outperforms full-parameter SFT and LoRA trained on the same data. The gain therefore comes from how the structural prior is represented and injected, not simply from in-domain supervision. SLP thus establishes a lightweight, parameter-efficient paradigm for human-centric spatial perception in MLLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.