Harnessing Multimodal Large Language Models for Training-free Human-object Interaction Detection
Abstract
Human-object interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Traditional supervised methods perform strongly but rely on task-specific training. Recent multimodal large language models (MLLMs) offer a promising route to training-free HOI detection through their broad visual-semantic knowledge. However, existing approaches leave semantic reasoning and visual perception loosely coordinated, inducing severe perceptual passivity where early semantic assumptions bias relation grounding and lead to self-reinforcing semantic circularity. To address these limitations, we propose HarnessHOI, a training-free framework that transforms passive MLLM interpretation into active interaction-guided perceptual reasoning harness. Central to HarnessHOI is an interaction-guided perception mechanism that proactively projects emerging interaction hypotheses back into visual space, actively discovering instances and refining ambiguous evidence through targeted observation. Subsequent relation-agnostic geometric adjudication reconciles these multi-source hypotheses under persistent object identities, establishing a unified visual grounding that decouples final predictions from initial proposal biases. Extensive experiments on HICO-DET and V-COCO show that HarnessHOI establishes state-of-the-art results among training-free HOI methods, demonstrating the effectiveness of the proposed MLLM harnessing strategy for complex visual-structuring tasks. The code will be released upon the publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.