Reliability-Aware Test-time Reinforcement Learning for Anomalous Video Understanding
Abstract
Anomalous video understanding aims not only to identify abnormal events in videos, but also to interpret their objects, actions, interactions, and high-level semantics. Recent Video Large Language Models (Video-LLMs) have demonstrated promising zero-shot capabilities for this task, yet frozen models remain limited by diverse anomaly patterns and evolving deployment environments. Test-time reinforcement learning offers a promising solution by enabling models to adapt from self-generated feedback without additional human annotations. However, directly applying it to anomalous video understanding faces three reliability bottlenecks: (1) pseudo-labels may be unstable across semantically equivalent queries, introducing unreliable self-supervision; (2) coarse consensus rewards ignore generation uncertainty, weakening reward informativeness; and (3) unanimous rollout groups receive identical rewards, causing group-relative advantages to collapse and eliminating effective policy-gradient signals. To address these challenges, we propose a reliability-aware online test-time reinforcement learning framework that improves reliability at the sample, generation, and optimization levels. Specifically, dual-query consistency filtering retains samples with cross-query consistent pseudo-labels, an entropy-aware consensus reward combines answer agreement with generation uncertainty, and virtual-anchor advantage estimation preserves informative relative advantages for unanimous rollout groups. Experiments on VAU-Bench demonstrate consistent improvements over frozen and supervised baselines. Notably, on the ECVA subset with thinking, our method improves accuracy from 75.81% to 90.18% over the frozen backbone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.