ProgAlign: Progressive Semantic Alignment with Visual Chain-of-Thought Reasoning for Open-Vocabulary HOI Detection
Abstract
Open-vocabulary human-object interaction (OV-HOI) detection is a fundamental task for human-centric scene understanding. Existing end-to-end HOI methods often suffer from insufficient fine-grained visual semantics mining and limited reasoning-process interpretability. To address these challenges, we propose ProgAlign, an end-to-end framework that integrates visual chain-of-thought reasoning with a progressive semantic alignment mechanism. Specifically, a visual chain-of-thought module is designed to construct a multi-step coarse-to-fine HOI evidence chain across spatial, contact, and interaction stages, enabling the model to progressively capture fine-grained interaction cues, which in turn helps improve the discrimination of visually similar interactions. Furthermore, a progressive semantic alignment mechanism is introduced to align Large Language Model-derived stage-specific semantic priors with intermediate visual representations, thereby helping alleviate black-box reasoning issues and enhancing the semantic consistency of the intermediate reasoning process. Extensive experiments on the SWIG-HOI and HICO-DET datasets demonstrate that ProgAlign achieves state-of-the-art performance among comparable open-vocabulary methods, validating its effectiveness in advancing OV-HOI understanding and distinguishing complex and fine-grained interactive relations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.