Heimdall: Hindsight Distillation for Proactive Warning in Streaming Video
Abstract
Streaming vision-language models can judge whether a scene is dangerous. Proactive warning asks them to say so before the accident, within a preset cap on false alarms, and on the device beside the camera. However, training such a warner requires supervision, and the annotated accident times that anticipation methods train on are expensive to obtain and absent for a newly deployed camera's footage. We observe that a frozen large verifier is uncertain at the moment a warning is due but confident a few seconds later, once the accident is in view. This later verdict can become a training target for the earlier moment without any annotation. We propose Heimdall, which queries a frozen 8B verifier a few seconds after each window of the training stream, converts those later verdicts into a multi-horizon risk target for the window itself, and distills on that target a 2B model that runs in real time on a Jetson. At deployment the two models run as one online policy. Each alarms when its risk score crosses its threshold, and the two thresholds are chosen together so that the policy as a whole stays within the cap. On the held-out test split of PaSBench-Video, Heimdall warns before the accident on 1.5 times as many clips as a zero-shot model with the same verifier at half its false-alarm rate, and on a third more than a model distilled from the same verifier at the same moment. Published warners that find more accidents raise two to four times as many false alarms as the cap allows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.