Reasoning with Future Evidence for Streaming Audio–Visual Understanding
Abstract
Streaming audio–visual models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If that interpretation enters memory as a fact, later reasoning can keep repeating it even after audio contradicts it. We call this failure premature cross-modal commitment. We propose StreamFE, which expresses unresolved interpretations as claims about future evidence, linked to a verification interval, a verifying modality, and the states that depend on them. When the interval closes, StreamFE checks the claim against evidence from the specified modality. A refutation reduces the influence of the claim and its dependent states, then guides a state update using the new evidence. Separate audio and visual retention preserves the evidence needed for these checks, and an answer gate delays responses while answer-critical claims await review. Using a frozen Qwen3-Omni-30B-A3B-Instruct backbone with lightweight adaptation, StreamFE outperforms the strongest open baselines on five streaming and audio–visual benchmarks by more than 10% relative on average. We also introduce AudioFaith, which holds video fixed and edits audio to test agreement, absence, contradiction, coexistence, and subtitle–speech conflict. StreamFE reaches , compared with at most for open baselines, while reducing vision-induced auditory hallucinations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.