acceptodds
Under review as a conference paper at ICLR 2027

Everything a Correct Answer Never Told You: Learning from Expert-Corrected Explanations in Colonoscopy Video QA

Abstract

In colonoscopy video question answering, a correct answer can rest on the wrong visual evidence: across the general-purpose open-source models we evaluate, 10.60% of correct responses are justified by hallucinated observations. Motivated by this finding, we introduce ColonExplainVQA, a dataset for evaluating multimodal models and for training them to answer from corrected visual evidence. It contains 7,962 video-question pairs across nine tasks from three public datasets. We minimally correct hallucinated claims in conclusions generated by Qwen 3.7 Plus while preserving supported content and retaining the original responses. These corrections provide supervision for learning both the reference answer and the visual evidence that supports it. Using ColonExplainVQA, we develop Qwen-Colon-8B by fine-tuning Qwen3-VL-8B-Instruct through supervised fine-tuning followed by direct preference optimization in modules selected through Colon Function Layer (CFL) analysis. On our manually reviewed test set, Qwen-Colon-8B leads all fifteen baselines, including proprietary models far larger than itself, with 79.15% answer accuracy; among its correct answers, the hallucination rate falls to 0.95%. Answer accuracy alone is an incomplete target; the evidence behind an answer can be curated, measured, and improved. Upon acceptance, we will publicly release the ColonExplainVQA annotations, clip-reconstruction metadata, and all associated code, subject to source-data access and licensing terms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.