VERITASMATH: Scoped Evidence and Bounded Authority for LLM Mathematical Reasoning
Abstract
Generated verifiers for large language model (LLM) mathematical reasoning can execute successfully while checking the wrong claim, omitting constraints, or covering only a finite subset of a universal statement. Reliability therefore depends not only on whether verification occurs, but also on what each piece of evidence is authorized to conclude. We introduce , an evidence-constrained reasoning architecture that operationalizes bounded epistemic authority through claim-level evidence contracts and an authority-constrained state machine. Each verification record captures the checked claim, polarity, domain, method, coverage, and provenance; executable conflict re-computation, selection-only arbitration, repair re-entry, and answer-locked rendering constrain how evidence can affect candidate selection and output, while an obligation graph tracks whether all requested deliverables are supported. On a pre-specified, latency-capped six-benchmark subset spanning final-answer and proof tasks, achieves a problem-weighted mixed aggregate, compared with for code-based self-verification, a -point improvement. In fixed-trajectory replays, the same contracts are associated with lower unsupported promotion and payload mutation, and with higher deliverable completeness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.