CG-OAS: ORDER-CONSISTENCY-GATED ANSWER SELECTION FOR MEDICAL RETRIEVAL-AUGMENTED GENERATION
Abstract
Medical RAG suffers from evidence-order sensitivity — the same set of evidence can produce inconsistent answers simply due to different ordering of the input. While prior work typically treats this as a robustness weakness, we find that this order sensitivity itself can be repurposed as a usable reliability signal. Through evidence-order perturbation (generating variants by permuting the same evidence set), examples fall into two regimes: strict-consistency (all permutations converge to the same answer) and non-strict (disagreement across permutations). The non-strict regime directly pinpoints where the model's evidence comprehension is fragile, providing a principled basis for answer selection. Motivated by this finding, we propose CG-OAS, a consistency-gated answer selection protocol. CG-OAS preserves the predicted answer when all evaluated evidence permutations agree; otherwise, it applies a lightweight, dataset-specific selector trained on non-unanimous validation examples. The selector uses only features derived from hard predictions, requiring neither evidence text nor model probabilities. The protocol returns an answer for every question and operates on outputs from validation-selected LoRA checkpoints without modifying retrieval or requiring additional generator training. On the MedMCQA and MedQA evaluation sets, learned selection improves accuracy over majority voting on non-unanimous questions by 2.90 and 4.98 percentage points, respectively, while leaving unanimous predictions unchanged. Overall, CG-OAS achieves a macro-averaged accuracy of 61.14%, exceeding original-order and majority-vote LoRA RAG by 1.36 and 1.04 percentage points, respectively, with paired bootstrap confidence intervals supporting both macro-averaged gains. These results show that evidence-order perturbations can provide actionable signals for targeted answer selection in medical RAG without updating the generator.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.