MS-ADD: MLLM-Supervised Asynchronous Dual-Branch Diffusion for Temporal Action Segmentation
Abstract
Weak supervision makes temporal action segmentation cheaper to annotate than frame-wise labels, yet every format of it still requires a person to watch and annotate every video. We use a multimodal large language model (MLLM) instead of a person to generate the supervision, which eliminates the human labour cost. Because the labels an MLLM generates inevitably contain errors, we select sparse timestamps and action sets, two formats of supervision that are easy for the MLLM to generate and that tolerate its errors. Neither format is perfect, so we use both: when one contains an error, the other may still provide valid supervision. The two formats differ in the computation they require, so under a fixed budget an allocation plan decides which share of the videos receives each format. Training on such supervision poses two challenges: a format may be missing from a video, and on average the timestamps of a video cover less than 1% of its frames. We propose MLLM-Supervised Asynchronous Dual-Branch Diffusion (MS-ADD). Its two-branch architecture takes the two formats of supervision as separate inputs and lets its branches exchange information. Its asynchronous diffusion assigns an independent noise level to every frame and every action, and thereby unifies pseudo-label generation, training and inference with the same forward pass. On CrossTask, Breakfast and GTEA, our model outperforms each baseline on nearly every segmental metric under the same single format of supervision, and a suitable allocation enables it to outperform the single-format state of the art at a lower budget. These results yield a general allocation plan.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.