EviPair: Paired Behavioral Diagnostics for Evidence-Responsive Decisions in Language Models
Abstract
Evidence-based decisions should follow changed premises and retain uncertainty when decisive information is missing. We introduce EviPair, a paired behavioral diagnostic combining entailed hints, premise reversals, irrelevant additions, and verified evidence removal on the same underlying problems. Joint correctness tests whether a model follows both premise states; a shared unknown condition tests whether it avoids unsupported assertions. This design measures Evidential Myopia as sensitivity to making an already entailed consequence explicit, while separating hint benefit from distraction. Our relational, numerical-aggregation, and natural-language studies cover two open-weight backbones and multiple hosted models, including a seven-family API extension. A screened Qwen2.5 configuration exhibits a 7.42-point hint benefit and a 7.03-point distraction cost in joint correctness. Hosted configurations show varied response profiles, including near-ceiling screened performance. Five-view SFT checkpoints improve premise tracking and uncertainty handling over unadapted backbones on new backgrounds from the same synthetic construction, without test-time gold hints. Matched five- view comparisons find no stable additional benefit from entailment consistency over SFT or paraphrase controls. Together, these results provide a paired evalua- tion protocol for distinguishing useful evidence interventions from distraction and unsupported certainty.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.