Where to Label and How to Adapt at Test Time for Temporal Action Segmentation
Abstract
This paper studies where to request frame-level annotations and how to use them to adapt a pretrained temporal action segmentation model at test time. To determine where to label, we propose an uncertainty-guided selection strategy that partitions each test video into consecutive regions with approximately equal uncertainty mass and queries one representative frame from each region. This design allocates annotations more densely to uncertain portions while retaining broad temporal coverage. To determine how to adapt, we use the queried labels as sparse semantic supervision and exploit their agreement with model predictions to dynamically regulate the adaptation trajectory across epochs. Experiments on standard benchmarks show consistent improvements under highly sparse annotation budgets. On Breakfast with MS-TCN, our approach improves accuracy from 65.4% to 81.9%, while querying only 0.16% of video frames on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.