Question-Guided Self-Distillation for Reducing False-Positive Hallucination in Vision-Language Models
Abstract
Vision-language models can affirm nonexistent objects, incorrect attributes, and invalid relations, or answer questions whose visual premises are false. Reduc- ing these false positives requires correcting unsupported content while retaining valid answers. We investigate question-guided self-distillation, which uses fac- tual conditions implicit in the original question as training guidance. Verification subquestions are supplied to a frozen base model without subquestion answers. A student generates responses from the original image and question and learns from the conditioned teacher’s token distributions on those same response prefixes. A second distribution-matching term constrains departure from the base model un- der the original input. Both terms use token-level Jensen–Shannon divergence; deployment requires neither question decomposition nor additional verification calls. With Qwen3-VL-8B and a 60K training pool, wh paired accuracy improves over our FINER-DPO baseline from 46.62% to 57.99% on FINER-CompreCap and from 36.42% to 47.01% on FINER-DOCCI. On DOCCI wh queries, negative- query accuracy rises from 46.09% to 61.03%, while positive-query accuracy falls from 83.60% to 80.63%. Other query settings and general capability evaluations reveal further trade-offs. These results show improved direct false-premise han- dling with the complete training recipe, while identifying selective correction and capability retention as remaining challenges.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.