acceptodds
Under review as a conference paper at ICLR 2027

Seeing the Evidence Is Not Enough: Removal Targets in Harmful Video Understanding

Abstract

Harmful video understanding increasingly pairs video-level judgments with temporal annotations of where harmful content occurs, and removing the annotated content offers a natural test of whether a model relies on it. Harmful videos make this test fragile: a slur, threat, or hateful message can recur in speech, on-screen text, and later scenes, so an annotation can mark valid evidence while harmful support persists elsewhere in the model's observation. We distinguish the *intervention target*, the information a deletion withholds, from the *claim target*, the support to which the resulting prediction change is interpreted, and call their correspondence *removal fidelity*. We study it with **HarmfulFaith** on three harmful-video datasets, four models, and sampled-frame and ASR/OCR text observations. Removing the annotated content changes far fewer `HARMFUL` predictions than retaining it preserves, and a clear harmful cue survives annotation-exact removal in 43–54% of audited observations. Source annotations cover a median 69.2% of independently audited text support and fully cover it in only 18.2% of cases with audited support; under the same 16-frame construction, visual harmful support is fully covered in 6.0% of such cases, versus 38.2% for Charades-STA event support. This residual support matters: source removal changes 13.3–19.5% of text predictions when a clear residual cue remains, versus 95.7–100% when none does, and the gap to audited-support removal widens as coverage falls. Under the same budget, removing audited support changes 93.8% and 88.9% of DeepSeek and GPT predictions, against 34.2% and 31.7% for matched random deletion. A weak response to removing annotated harmful content therefore need not mean that a model ignores harmful content: deletion-based conclusions about harmful-video models should be conditioned on what support the intervention actually withholds.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.