RethinkFuse: Cross-Window Evidence for Micro-Expression Event Decoding
Abstract
Long-video micro-expression analysis requires an event decoder to determine whether a local response is genuine and which temporal boundary best represents it. A single-window decoder couples candidate survival and boundary construction to one temporal scale, leaving no independent evidence for either decision. We introduce RethinkFuse, an event decoder for frozen joint spotting-recognition backbones. Its central observation is that cross-window agreement provides evidence of event support, while boundary disagreement exposes complementary localization hypotheses. RethinkFuse decodes the same score sequence with micro and macro windows, associates compatible candidates into event clusters, constructs a stronger single-window candidate and a fused candidate , and uses an Adaptive Gated Fusion router to reject the cluster, keep , or select . The router is fitted using only training subjects in each outer fold and leaves backbone parameters and frame-level predictions unchanged. Under fully nested subject-disjoint evaluation, RethinkFuse improves the spotting-recognition synergy score (STRS) by approximately 9.6% on SAMMLV and 30.1% on CAS(ME) relative to native decoding. Controlled comparisons attribute a positive STRS gain to adaptive routing on SAMMLV. On CAS(ME), the corresponding interval reaches zero at the reported precision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.