When Rewards Break Traceability: Training Traceable Evidence Paths with LedgerRL
Abstract
In high-stakes settings such as medicine and finance, a reader must be able to check the evidence behind an answer: it should be correct, reach the evidence it needs, and expose small units that can be checked. Retrieval-augmented reasoners are usually trained with outcome or citation rewards and judged by accuracy or citation precision; we find that such training admits four shortcuts that break these requirements: outcome rewards leave answers without a cited path, entailment rewards encourage claims that copy their evidence, citation rewards attach many sources to broad assertions, and formats that answer in one context search little. We answer each with a matching design. A typed claim–evidence ledger lets named claims cite evidence and the answer cite claims; staged execution begins with an explicit investigation; and LedgerRL trains one shared policy with claim-level rewards that discount copied claims and count only answer-used claims toward quantity credit. On three multi-hop corpora after medical-only training, the staged ledger elicits the search that a flat citation prompt skips, and even untrained it makes more answers traceably correct than a trained citation-RL policy. Unguarded claim credit collapses traceability, whereas the guarded rewards reach 19.7% complete-path correctness, above every baseline under the same prompts, and leave fewer answers uncited than an outcome-only ledger. No arm dominates: against citation RL and a citation cap that searches about as often, LedgerRL leads at claim-level TC@2 and covers more gold evidence than the cap, but not at final-answer check units, in gold precision, or in gold-citation F1; it leaves more answers uncited and trails the cap by 7.9 points in accuracy at 1.9 times its tokens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.