Embracing Ambiguity: Hypothesis-Guided Linguistic Policy Optimization for Video Anomaly Reasoning
Abstract
Video anomaly understanding requires models to identify anomalous events and explain why they are anomalous. Existing approaches improve recognition through informative observations or reusable linguistic guidance, yet a correct prediction may still rely on a criterion that misclassifies normal events. We introduce Anom-Ψ, a hypothesis-guided linguistic policy optimization framework that turns differences between competing interpretations into feedback for refining anomaly criteria while keeping model weights frozen. During training, the model interprets shared observations around temporal changes under normal and abnormal hypotheses, exposing differences in the evidence and contextual assumptions behind its judgments. A linguistic optimizer edits a hierarchical experience bank using reward-guided comparisons, keeping hypothesis-derived corrective references distinct from neutral-trajectory preferences. We derive a sufficient condition for these linguistic policy updates to improve expected reward under weak supervision. Across benchmarks, Anom-Ψ outperforms the compared tuning-free baselines, reaching % Micro-AUC on MSAD with 20% of its training data. Ablations and repeated runs show that opposing-hypothesis comparisons complement neutral preferences, yielding more effective and consistent experience learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.