QRePair: Question-Guided Reconstruction via Paired Prototype Retrieval for Missing-Modality Audio-Visual Question Answering
Abstract
Audio-Visual Question Answering (AVQA) requires reasoning over audio, visual, and linguistic cues, yet one modality may be unavailable at inference time, leaving question-relevant evidence unobserved. Semantic retrieval can provide useful priors, but may not preserve structured evidence such as source count or temporal configuration. We propose Question-Guided Reconstruction via Paired Prototype Retrieval (QRePair), which reconstructs task-relevant missing-modality representations from paired temporal prototypes. Because relevant evidence depends on the question, Question-Guided Modality Encoding first produces question-conditioned audio and visual representations. Paired Prototype Construction organizes complete training representations into temporally resolved audio-visual prototype pairs, and Paired Prototype Retrieval uses the available representation to retrieve the corresponding missing-modality prior. Question-Guided Reconstruction then adapts this prior to the current question and the feature space of a frozen answering backbone. On 5,096 audio-visual test questions from MUSIC-AVQA, QRePair achieves 71.59% and 73.49% accuracy with visual and audio inputs missing, respectively, representing gains of 2.22 and 1.43 percentage points over the reported R^2ScP results. Ablations further show that disrupting audio-visual correspondence reduces average missing-modality accuracy by 4.25 points, while removing explicit question conditioning from reconstruction reduces it by 8.96 points. QRePair also shows strong performance on AVQA and fully observed MUSIC-AVQA-R and MUSIC-AVQA 2.0 benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.