Keep or Repair? Certified Acceptance for Conservative Query Repair
Abstract
Language-model repairers are a natural add-on to knowledge base question answering (KBQA) parsers, but once the parser is accurate, repair is governed by the damage it does rather than the errors it fixes. When upstream accuracy reaches 92%, as on our KQA Pro development cohort, even a small fraction of wrongly edited correct answers can outweigh gains on incorrect questions. We formalize conservative query repair and separate what to propose from what to accept. Our certified acceptance gate accepts a proposed query only when the repairer's own KEEP-sequence probability κ on it exceeds a threshold locked by a Clopper–Pearson bound; the rule provably controls the population regression rate, needs no training, and applies to repairers that expose KEEP token probabilities. On 11,797 KQA Pro questions, the gate turns a net-negative frozen repairer (net −18) into net +57, with a 95% upper bound of 0.35% on its regression rate against a 0.5% target, and κ achieves higher proposal-ranking AUROC than token likelihood or a prompted judge. Fine-tuned repairers achieve similar net gains (+55 to +57), raising KEEP recall from 71% to 88% without improving error-ranking AUROC. After target-domain recalibration on WebQSP and CWQ, the fine-tuned repairer's gate never exceeds its risk target across 1,000 random splits, but yields no positive net gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.