acceptodds
Under review as a conference paper at ICLR 2027

MMA-MTU: Benchmarking Multi-Task Micro-Action Understanding

Abstract

Multimodal large language models (MLLMs) have shown strong capabilities in video understanding, yet existing benchmarks largely focus on daily activities and general actions. Their ability to perceive and understand micro-actions, which are brief, low-amplitude, and spatially localized, remains insufficiently assessed. Beyond recognizing action categories, micro-action understanding requires reasoning about temporal relations, locating action boundaries, recovering relevant events, and describing how movements unfold. Our work addresses this evaluation gap through five key contributions: **1) Multi-task benchmark.** We introduce **MMA-MTU**, a benchmark that jointly evaluates Recognition, Temporal Reasoning, Grounding, Detection, and Description on shared video evidence. **2) Large-scale, multi-domain data.** MMA-MTU contains 4 domains, 11,180 untrimmed videos, and 122,845 queries covering 82 fine-grained categories, and 5 tasks with 12 subtasks. It provides dedicated training, validation, and test splits for instruction tuning and model learning. **3) Paired queries and fine-grained descriptions.** MMA-MTU includes 34,876 label-free queries paired with label-based counterparts, replacing action names with descriptions of observable movements. It also provides motion-evidence-guided event descriptions that capture how individual actions are performed. **4) Extensive evaluation.** We evaluate **36** MLLMs under the same protocol. The results expose a clear divide between different tasks: the best Temporal Reasoning accuracy reaches 66.62%, whereas the best Grounding average R@1 and Detection average mAP reach only 26.65% and 5.44%, respectively, showing that success on one ability does not reliably transfer to the others. Supervised fine-tuning significantly improves performance across all 5 tasks. **5) Downstream exploration.** A preliminary emotion-recognition study further explores how predicted micro-action information can support human emotion understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.