acceptodds
Under review as a conference paper at ICLR 2027

Learning to Retrieve Multimodal Evidence from the Answer Alone: Latent Evidence-Set Retrieval for Multimodal Multi-Hop QA

Abstract

In multimodal multi-hop question answering, a retriever selects images, passages, and tables from a question's candidate pool as evidence for a vision-language model (VLM) reader. Training such retrieval-augmented systems typically requires evidence annotations alongside questions and answers, but these annotations are costly and scarce. Learning retrieval from answers alone is difficult: answer quality does not identify which sources were useful or which work together, and pretrained retrievers favor particular modalities regardless of what the reader can use. We introduce LEAF, a two-stage framework for building multimodal RAG systems without supporting-source supervision. First, we fine-tune the reader on reference answers and retrieved contexts of varying sizes and diversity, creating a task-adapted answer generator and teacher. Second, we freeze it and use answer rewards to adapt the retriever encoder and to train a selector that chooses evidence diversity and size per question. Across Qwen and VLM2Vec retrievers and Qwen and Gemma readers, LEAF outperforms the two standard annotation-free policies, Top-K and fixed MMR, on MultiModalQA and WebQA, reaching 76.02 F1 and 53.75 Overall QA respectively. These scores exceed published evidence-supervised SKURG and Solar results on both benchmarks and MuRAG on WebQA, and our best WebQA system ranks fifth all-time by Overall QA on the official leaderboard.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.