acceptodds
Under review as a conference paper at ICLR 2027

Multiple Questions, Shared Evidence: Budget-Aware Grouped Inference for Long-Video QA

Abstract

Long videos are frequently asked about using multiple questions instead of a single one, and a fixed visual-token budget must cover the entire batch; naively re-encoding video content for every question wastes this budget as the batch grows, turning batch video question answering into a resource-allocation problem. Existing methods sit at two extremes: answering each question with its own context repeatedly re-encodes overlapping evidence, while answering the whole batch with one shared context discards question-specific details, and neither adapts to how much the questions actually overlap. We introduce QShare, a batch VideoQA framework that treats which questions should share a context and how the shared budget is spent as one joint decision rather than committing to either extreme. QShare forms an initial partition from question–frame evidence overlap, allocates frames across groups with a submodular greedy procedure that provably approximates the optimal coverage under the shared budget, and iteratively refines the partition — moving, splitting, or merging questions — based on the contexts produced by the allocation; it then answers each group with a single multimodal language-model call. Across CinePile, CG-Bench, and LVBench with two backbones, QShare improves average accuracy by 4.0 points over the strongest baseline under matched batch budgets, reduces the number of LMM calls and the amount of repeated visual prefilling, and keeps its accuracy advantage as the number of questions per video grows.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.