acceptodds
Under review as a conference paper at ICLR 2027

ProgAlign: Progressive Semantic Alignment with Visual Chain-of-Thought Reasoning for Open-Vocabulary HOI Detection

Abstract

Open-vocabulary human-object interaction (OV-HOI) detection is a fundamental task for human-centric scene understanding. Existing end-to-end HOI methods often suffer from insufficient fine-grained visual semantics mining and limited reasoning-process interpretability. To address these challenges, we propose ProgAlign, an end-to-end framework that integrates visual chain-of-thought reasoning with a progressive semantic alignment mechanism. Specifically, a visual chain-of-thought module is designed to construct a multi-step coarse-to-fine HOI evidence chain across spatial, contact, and interaction stages, enabling the model to progressively capture fine-grained interaction cues, which in turn helps improve the discrimination of visually similar interactions. Furthermore, a progressive semantic alignment mechanism is introduced to align Large Language Model-derived stage-specific semantic priors with intermediate visual representations, thereby helping alleviate black-box reasoning issues and enhancing the semantic consistency of the intermediate reasoning process. Extensive experiments on the SWIG-HOI and HICO-DET datasets demonstrate that ProgAlign achieves state-of-the-art performance among comparable open-vocabulary methods, validating its effectiveness in advancing OV-HOI understanding and distinguishing complex and fine-grained interactive relations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.