Compositional Evaluation of Human-Object Interaction Understanding with Real and Synthetic Multi-Tag Scenes
Abstract
Human–object interaction (HOI) understanding requires models to distinguish multiple people, objects, and interactions that may coexist within the same scene. However, existing HOI evaluation commonly reports aggregate performance or analyzes isolated interaction configurations, providing limited insight into model behavior under compositional scenarios involving multiple sources of ambiguity. We first analyze multi-tag images from HICO-DET and evaluate HOI detection models on combinations of interaction configurations. Our analysis reveals that real multi-tag images are limited and highly imbalanced across combinations, making systematic compositional evaluation difficult. To address this limitation, we construct an image-conditioned synthetic evaluation set by editing real HICO-DET images to introduce controlled combinations of human–object interactions while preserving the original scene context, lighting, and visual complexity. Because the intended interaction configurations and HOI classes are specified during generation, the corresponding labels are obtained automatically and verified through human quality control. Using the real and synthetic subsets, we evaluate both specialized HOI detection methods and vision-language foundation models under a unified compositional framework. Our study provides a systematic analysis of how multiple interaction factors jointly affect model behavior and offers a scalable approach for evaluating compositional HOI understanding beyond the limited coverage of existing benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.