QMA-RAG: Query-Conditioned Multi-Granularity Context Construction for Long-Video Question Answering
Abstract
Multi-granularity retrieval offers local observations and broader context for long-video question answering, but several retrieved nodes may refer to the same source frames. Constructing a limited visual context therefore requires deciding how these semantic views jointly affect which observations are selected. We introduce QMA-RAG, which combines multi-granularity support on shared source clips before selecting the visual context. A reusable graph links atomic, event, and macro descriptions to their supporting clips. Question-conditioned retrieval supplies candidates whose scores are fused on these clips; selection balances the resulting support with temporal needs and repeated coverage. Selected clips are converted into a bounded set of unique frames with original temporal information and retrieved structured support. Across MLVU, Video-MME, and LongVideoBench with three answering backbones, the complete system improves Uniform inference in all nine settings by 3.80 percentage points on average. On MLVU, construction from identical candidates improves over exact frame deduplication by 1.66 points. Removing cross-level fusion while retaining the full selector reduces accuracy by 1.93 points, supporting the role of aggregated semantic support in source selection. Answering-input and component controls examine how the resulting context and its supporting design choices contribute to answering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.