GTR-HOI: Ground, Transport, and Reinforce for Weakly Supervised HOI Detection
Abstract
Weakly supervised human–object interaction (HOI) detection learns to localize and recognize interactions from image-level labels, without explicit correspondences between labels and human–object pairs. This ambiguity can reinforce incorrect interaction assignments during training. We propose GTR-HOI, a three-stage framework that grounds image-level labels, transports localization supervision, and reinforces interaction predictions. First, a pretrained multimodal large language model (MLLM) generates pseudo-ground truths from human, object, and relation perspectives, with spatial agreement determining their confidence tiers. To account for grounding uncertainty, unbalanced optimal transport transfers localization supervision to detector pairs through soft correspondences. Pair expert routing integrates appearance, spatial, and contextual cues for pair validity prediction and visual HOI scoring. Hierarchical supervision uses consensus assignments with high or medium confidence to train the visual scoring branch and fine-tune the MLLM for semantic HOI scoring. Finally, cross-prompt reinforcement learning uses semantic feedback across four prompt views to refine decisions to accept or reject candidate interactions through weighted proximal policy optimization. Spatial and fixed reference certainty weights regulate these updates to limit the influence of unreliable supervision. Experiments on HICO-DET and V-COCO demonstrate state-of-the-art performance among the compared weakly supervised methods, achieving 36.44 Full mAP and 60.6 Scenario 2 role AP, respectively. The code is provided in the supplementary materials.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.