ForeWarn: Risk-Adaptive Inference for Proactive Safety Warning in Streaming Video
Abstract
Streaming video models now describe what is happening in front of a camera, and proactive safety warning asks them to act on what is about to happen. However, existing warning methods decide from the current state, and confirming a hazard once it is visible consumes part of the time left to warn. Our diagnostics reveal that serial confirmation filters out early candidates and that a later confirmation recognizes more but returns too late to warn. In this paper, we propose ForeWarn, which lets the near-future risk predicted by a small vision-language model decide both when to warn and when to verify, so that a forecast acts before confirmation where time is short and waits for it where time allows. The same risk read at two horizons drives an immediate path that does not wait for a verifier and an advisory path that combines predicted risk with a returned state judgment; a risk-dependent schedule decides when to query the verifier. Furthermore, calibration on normal footage sets the advisory thresholds per deployment, and a one-step analysis ties the value of verification to confirmation quality and timely return. Extensive experiments on five public domains, self-recorded live trials, and three Jetson edge devices demonstrate that ForeWarn raises macro pre-incident recall to 1.4 that of the strongest of four recent baselines at similar verifier cost, with gains of up to 47.7 points on a single domain and up to 110 less verifier time per clip.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.