acceptodds
Under review as a conference paper at ICLR 2027

LongHOI-Bench: Benchmarking and Improving Non-Contact Human-Object Interaction Understanding in MLLMs

Abstract

Understanding non-contact human-object interactions requires identifying a person's target from gaze, pointing, or contextual evidence. While multimodal large language models (MLLMs) support diverse visual tasks, existing benchmarks largely evaluate these cues separately, limiting systematic assessment of person-target understanding. To address this gap, we introduce **LongHOI-Bench**, containing 6,240 images and 56,880 questions spanning localization, attribute discrimination, spatial relations, and contextual intention. Human-verified targets and typed negative alternatives connect task performance with three grounding errors: *drift*, *distractor*, and *ambiguity*. We benchmark 13 proprietary and open-source MLLMs on LongHOI-Bench. Results reveal substantial limitations, with the strongest model reaching 68.5% mean accuracy across interaction types while humans exceed 92% in each interaction category. A balanced error audit identifies distractors as the most frequent failure. Motivated by these findings, we propose **CLASP**, which combines *error-type diagnosis* with *contrast-based layer selection* and weighted suppression through analytic weight editing. On paired localization queries, CLASP corrects 69.8% of distractor errors while introducing errors in 4.2% of initially correct predictions. Direction controls show that diagnostic relevance matters beyond captured contrast energy; the intervention also improves external gaze and pointing performance without recalibration. The benchmark and code will be publicly released upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.