acceptodds
Under review as a conference paper at ICLR 2027

Structured Provenance and Inference Boundaries for Evidence-Grounded Political Reasoning in Language Models

Abstract

Language models can produce fluent, evidence-grounded answers while still exceeding what the underlying evidence actually warrants—for example, converting temporal sequence into causation, actor-specific statements into institution-wide claims, or limited archival access into claims of historical absence. We investigate whether structured provenance information and explicit typed inference-boundary instructions improve a model’s agreement with the maximum claim licensed by bounded political and institutional evidence. We introduce a trace-locked evaluation framework built from 24 items nested within 12 independent evidence traces drawn from authoritative U.S. government and institutional materials. Three language-model families each completed 96 judgments across four conditions: bounded evidence alone; evidence plus structured provenance; evidence plus provenance and typed inference-boundary controls; and a provenance-matched sham condition controlling for added instruction length and salience. The primary outcome is trace-weighted exact agreement with a frozen procedural reference across claim warrant, evidence state, relation warrant, and scope. After an exact deterministic correction for singleton-array serialization errors, the trace-weighted boundary-minus-evidence difference was 0.083 for each of the three model families. Exact paired tests yielded p ≥ .5, and all descriptive bootstrap intervals included zero. We therefore find no statistically persuasive evidence, at this feasibility scale, that provenance or typed boundary prompting improves reference agreement. However, the evaluation exposed a distinct methodological failure mode: 47 of 288 original outputs violated the frozen schema, and a planned model-based “format-only” retry changed substantive judgments in all affected cases, with failures concentrated in the boundary condition. These results suggest that structured-output evaluations can confound semantic reasoning quality with interface compliance and repair behavior. The principal contribution is therefore methodological: evidence-grounded language-model evaluation should separately measure semantic judgment, schema compliance, uncertainty and scope preservation, and any post-generation repair process. More broadly, the study provides a reproducible framework for testing whether language models remain within explicitly defined evidentiary and inferential limits without treating procedural reference agreement as political correctness, objective historical truth, or proof of general reasoning competence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.