More Correct Labels Can Hurt: Delayed Feedback in Streaming Activity Recognition
Abstract
In streaming activity recognition, a requested label may arrive after the activity has changed. Unsupervised test-time adaptation updates batch statistics (BN), minimizes entropy (Tent), or filters predictions (OATTA), while delayed replay (IWMS) learns from historical labels. These approaches leave open how a returned label should enter the current activity state and influence subsequent predictions. By analyzing errors at label arrival, we find that directly injecting a correct historical label can overwrite a correct current prediction. Local loss analysis further shows that even a useful correction can become harmful when applied too strongly. We therefore propose Gated Arrival, which separates historical learning, state injection, and output mixing. A calibrator learns from returned labels and their saved origin predictions. Two gates fitted offline use label age and current predictions to set the weight of state injection and its later contribution to the output, with the recognition backbone frozen. Experiments on five activity datasets under simulated response delays show that additional correct labels can reduce recognition and that delay can reverse feedback-rule rankings. Gated Arrival reduces errors caused by labels that no longer match the current activity. Component comparisons show how historical calibration and state correction affect recognition and probability estimates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.