acceptodds
Under review as a conference paper at ICLR 2027

Self-calibrating Agentic Generation and Answering over Multimodal Instructional Collections

Abstract

Instructional recordings such as lectures, tutorials and talks convey information through speech, slides and source documents like research papers. Question an- swering (Q/A) over them therefore requires reasoning across modalities, combin- ing shared with modality-specific information. Existing instructional Q/A bench- marks do not test this together with demanding reasoning constraints such as tem- poral dependencies, multi-hop connections and long-context reasoning. We pro- pose MM-AdvQAGen, a multimodal Q/A generation pipeline that certifies each question’s modality, context and multi-hop requirements through a series of falsi- fiable tests run by blind solvers. It yields MMAv-InstructBench, built from the seed datasets M3AV (Chen et al., 2024) and SlideSpeech (Wang et al., 2024a): 4,624 questions (1,618 and 3,006) for evaluating answering across four dimen- sions, temporal grounding, cross-modal synthesis, long-context reasoning and multi-hop reasoning. For retrieval and answering over such collections we further propose MM-Graph-RAG, a graph-retrieval pipeline that builds one cross-modal graph per recording, linking slide text, speech transcript and related document content, and derives its retrieval settings from the recording itself rather than per- dataset tuning. On MMAv-InstructBench, MM-Graph-RAG improves judge accuracy over the strongest baseline by up to 10.7 percentage points on M3AV and 16.6 on SlideSpeech. On external datasets of other types it likewise wins against the best baselines by good margins.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.