Video-XAI Bench: A Known-Evidence Benchmark for Spatiotemporal Explanation
Abstract
Video classifiers flag events for human review in public-safety and security monitoring. A classifier's decision is hard to trust unless a reviewer can check the evidence behind it. Explanation methods are meant to provide this evidence by highlighting the regions and moments of a video that drove the decision. Existing evaluations check if explanations look plausible or if removing the highlighted content lowers the model's confidence; they do not test if explanations point to the right place and moment. We introduce Video-XAI Bench, a benchmark for explanations of video classifiers. It tests explanation methods in two settings. The first uses real violence-recognition videos. The second adds a moving pattern to real videos at a known place and time, so we know exactly where and when the evidence is. On real videos, no method consistently performs best, as the best choice changes with the dataset, the model, and the score. With known evidence, scores based on the model's confidence do not reveal if an explanation finds the right moment, and some of the explanations that find it best partly reflect the input video rather than what the model learned. In a pilot with 18 participants, those who agreed that a video showed aggression often disagreed about which frames showed it, so human-marked evidence on real videos forms a window with soft edges. Our study concludes that video explanations should be tested against known evidence, not only against the model's confidence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.