SemAlign-VQA: Semantic Alignment for Medical Visual Question Answering
Abstract
Medical visual question answering requires correct answers to persist across changes in question formulation. We study how paired training changes this outcome, separating correct-answer retention from answer agreement. SemAlign-VQA combines original-question answer supervision with SemLoss, a pull–push regularizer on projected pre-answer states, using LoRA and a projector bypassed at inference. Evaluation uses an original question and five fixed context wrappers on PathVQA or automatically screened rewrites on SLAKE-English and VQA-RAD. At seed 42, Moondream2 improves all-six correctness by 1.85 and 2.17 percentage points on PathVQA and SLAKE, with nearly unchanged original accuracy. Matched-question analysis locates these gains in retention, while outcome transitions distinguish sustained correct answers from stabilized errors. The training controls reveal a second distinction: direct rewrite supervision reaches 65.22% all-six correctness on SLAKE versus 57.30% with original-only alignment, and the summary-level answer-aware NCE control nearly matches the margin objective on PathVQA despite lower agreement. Effects also vary across backbones and training seeds. Together, the paired comparisons show where alignment changes correctness, how those changes differ from consistency gains, and why supervision-matched controls matter when choosing a training objective.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.