LLM-GUIDED DISTRIBUTION PROPOSAL REFINEMENT FOR WEAKLY SUPERVISED TEMPORAL ACTION LOCALIZATION
Abstract
Weakly supervised temporal action localization aims to localize action instances in untrimmed videos using only the video-level label. However, the semantic mismatch between weak supervision and background frames may yield inaccurate action scores, which in turn causes inaccurate action boundaries. To address this, we introduce an LLM-guided distribution proposal refinement model, which explicitly models subaction semantics via four mechanisms. First, to reduce the semantic mismatch in weak supervision, we design an LLM-guided proposal learning module. It uses LLM prompts to extract the textual semantics of start, middle, and end subactions. A proposal structure loss ensures that subactions are properly ordered. Second, to alleviate ambiguous prediction among subactions, we design a contrastive proposal learning module. The discriminativeness of subaction representations is optimized via intra-action and inter-action contrastive losses. Third, to alleviate ambiguous boundary localization caused by smooth transitions, we design a proposal transition learning module. The score changes at the boundary are encouraged to be separable via a boundary transition loss. Fourth, to alleviate ambiguous boundary localization among gradual transition segments, we propose a proposal distribution refinement module. The boundary is refined to the location by considering the difference between left- and right-window distributions (measured via the first-order Wasserstein distance). Experiments show that our method achieves state-of-the-art performance on THUMOS'14 and ActivityNet 1.3 datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.