Everything a Correct Answer Never Told You: Learning from Expert-Corrected Explanations in Colonoscopy Video QA
Abstract
In colonoscopy video question answering, a correct answer can rest on the wrong visual evidence: across the general-purpose open-source models we evaluate, 10.60% of correct responses are justified by hallucinated observations. Motivated by this finding, we introduce ColonExplainVQA, a dataset for evaluating multimodal models and for training them to answer from corrected visual evidence. It contains 7,962 video-question pairs across nine tasks from three public datasets. We minimally correct hallucinated claims in conclusions generated by Qwen 3.7 Plus while preserving supported content and retaining the original responses. These corrections provide supervision for learning both the reference answer and the visual evidence that supports it. Using ColonExplainVQA, we develop Qwen-Colon-8B by fine-tuning Qwen3-VL-8B-Instruct through supervised fine-tuning followed by direct preference optimization in modules selected through Colon Function Layer (CFL) analysis. On our manually reviewed test set, Qwen-Colon-8B leads all fifteen baselines, including proprietary models far larger than itself, with 79.15% answer accuracy; among its correct answers, the hallucination rate falls to 0.95%. Answer accuracy alone is an incomplete target; the evidence behind an answer can be curated, measured, and improved. Upon acceptance, we will publicly release the ColonExplainVQA annotations, clip-reconstruction metadata, and all associated code, subject to source-data access and licensing terms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.