acceptodds
Under review as a conference paper at ICLR 2027

Observed Accuracy Is Not Visual Capability: Auditing and Repairing Decision-Level Confounds in Video Anomaly Benchmarks

Abstract

A rule reading only the video header, meaning resolution, duration and frame count, predicts a video anomaly benchmark’s labels at 0.970. We had adopted that benchmark because it looked clean. How much of such a shortcut survives a strong model? We formalize the question as residual side-channel capacity, the drop in optimal predictive risk from adding an unintended channel to a frozen model score, which under log loss is I(Y ; Z | S); our plug-in estimate is a lower bound on it. Above the strongest model we measure, video headers close 42% of the label uncertainty its score leaves open, lifting AUC from 0.947 to 0.982. On a rendered benchmark the same audit returns capacity indistinguishable from zero, so this is a property of how footage was collected, not of the estimator. A benchmark score is an intermediate observable, and reading capability off it means first showing which layer moved. We therefore separate what one score conflates, acquisition shortcuts, visual ranking and the decision operating point, probing each by its own intervention and releasing the audit as a harness with cached scores that reproduces every headline number on CPU in seconds. Task structure makes part of any score unidentifiable, which gives an exact account of when repairing an operating point helps, hurts or does nothing. Because ranking transfers across domains and absolute scale does not, a certified rank-based selective rule retains 1.2–1.7× the coverage of an equally certified score-threshold rule at matched risk, with paired intervals excluding zero in three of four settings. We also report four revisions to our own conclusions and 18 pre-registered negative verdicts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.