Too Many to Judge? MultiDance: A Multi-Task Benchmark for Multi-Target Action Quality Assessment
Abstract
Action Quality Assessment (AQA) aims to quantify how well actions are performed, serving as a fine-grained video understanding task with broad applications. However, most studies evaluate a single athlete or team in isolation, overlooking competitions where multiple performers share the stage and interact. Such settings require target-specific assessment that accounts for competitive context and provides interpretable feedback. To address these issues, we introduce Multi-Target AQA (MT-AQA), a task that jointly evaluates and compares multiple competing targets within the same video. We present MultiDance, the first large-scale benchmark for MT-AQA, comprising 1,548 long videos from 99 events with 14,498 annotated performer instances of diverse skill levels and annotations for classification, ranking, scoring, and coaching. We further propose Multi-Target Assessment Coach (MultiAC), a multimodal large language model (MLLM) framework that unifies these tasks through domain knowledge prompts and coarse-to-fine dual-stage adaptation. MultiAC outperforms single-target AQA extensions and general MLLMs on MultiDance, where joint training reveals complementary benefits between single- and multi-target data. Results on EgoExo-Fitness further support coarse-to-fine multi-task adaptation beyond dance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.