CIViC: Evidence Gates for Audio-Conditioned Visual Prediction
Abstract
Removing sound from a visual predictor may change its output, but that test leaves the model's use of the intended audio information uncertain. We call this protocol Conditional Input and Visual Interpretation Checks (CIViC). It links a specific interpretation to the intervention, task comparison, or target check needed to support it. We apply the protocol to an actor-disjoint face-image task, a synthetic feature-mixture task, and a fresh source-video-disjoint SpeakerVid proxy task. The static face-image test detects a small response to replacing audio with another actor's; a separate temporal test does not establish an audio benefit on its endpoint. In a post-confirmation comparison on a second actor-disjoint raster target, four constant-audio fits score better than their audio-trained counterparts, despite a large response when audio is removed at test time. In the synthetic feature-mixture task, the added condition changes a 128-coordinate visual proxy but raises withheld-composition error from 0.1707 for audio alone to 0.1807 at 1,000 steps; the gap shrinks with longer fitting. In the source-video-disjoint SpeakerVid task, a same-target audio-zero arm has lower motion-proxy RMSE than aligned audio in all four held-out conditions and all five paired seeds on four accelerator platforms using one source pack. The same direction appears within each of the four held-out source-video groups, but a secondary long-form drift proxy favors aligned audio in all four conditions. Audio-only, dummy-condition, and shuffled-condition arms expose distinct endpoint behavior, showing why an output response alone cannot justify a conditional interpretation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.