acceptodds
Under review as a conference paper at ICLR 2027

Event-Aware Prototype Matching for 3D Human Motion Retrieval

Abstract

Text‑to‑motion retrieval is a challenging task that aims to search relevant motion sequences based on natural language descriptions. The typical solution is to directly align the global sequence‑level and sentence‑level features for text‑motion matching during learning. However, motion sequences often involve multiple action events, which are difficult to capture effectively with just a single embedding. In addition, different motion sequences may share similar actions, which leads to high semantic overlap between unpaired samples and thereby introduces semantic ambiguity during alignment. To this end, we propose an Event‑Aware Prototype Matching (EAPM) approach, which models text‑motion correspondence as an action events matching procedure. Specifically, to decompose events from motion sequence and text, we design an event‑aware prototype aggregation method that adaptively aggregates local features into a set of action event prototypes. We further design a variance loss to encourage these prototypes to focus on distinct contents, thereby effectively capturing diverse action semantics. Finally, we present a prototype‑based event soft‑alignment approach to modulate confidence estimation for events via intra‑modal self‑matching probability, which enables more fine‑grained text‑motion alignment at the event level. Extensive experiments show that our method significantly outperforms existing methods in text‑to‑motion retrieval and other challenging tasks, such as human interaction recognition and motion temporal localization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.