Trust at the margin: Provenance-visibility effects and model-dependent risk localization in LLM knowledge conflicts
Abstract
LLMs increasingly rely on external information that may conflict with their parametric knowledge. Recent work demonstrates that provenance can influence conflict resolution, however, three questions remain: how provenance visibility affects model decisions, where resulting risk is concentrated, and how this information can guide mitigation. In this paper, we address these questions in three steps. First, we introduce PROVMARGIN, a benchmark that varies provenance visibility while holding others fixed and measures model epistemic margins before intervention. Second, we characterize provenance effects and their risk localization. Across eight models, removing provenance cues indicating low credibility increases false answer adoption by 9.95%, with a larger effect near decision boundary. Extensions across model families, response measures, and tasks show that overall effect generalizes more consistently than its boundary localization. Real-world evaluations further reveal heterogeneous effects for erroneous claims, while official provenance consistently supports correct updates. Finally, we translate these findings into PROVROUTE, which combines provenance, epistemic margins, and observable responses to prioritize verification. At a budget of 1/6 of requests, PROVROUTE reduces residual false-adoption risk to 27.56%, improving over the strongest static baseline by 2.07 %. Together, these results connect causal measurement, model dependent risk localization, and selective verification, while emphasizing the need to reduce errors without hindering useful knowledge updates. and task requirements further moderate provenance sensitivity. A development only study of 30 real records finds heterogeneous effects for erroneous sources, while official provenance improves adoption of correct updates. We then propose PROVROUTE, which combines model-specific margins, source risk, and observable response signals to prioritize verification. On a separate outcome-unseen panel, reviewing one sixth of requests yields 27.56% residual false-adoption risk under an ideal-verifier assumption, improving on the strongest static margin–provenance baseline by 2.07 percentage points at 83.33% automatic coverage. These findings support preserving provenance and calibrating verification to the target model while retaining the utility of reliable external evidence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.