From Knowledge to Decisions: Repairing Evidence Seeking after Full-Evidence Adaptation
Abstract
Full-evidence adaptation can transfer non-uniformly across evidence-conditioned competence and interactive evidence acquisition. On HotpotQA, it raises Oracle exact match from 58.6% to 68.0% but lowers Interactive exact match from 40.6% to 32.0%. Restoring inherited Base acquisition while retaining the adapted reader raises Interactive EM to 46.1%, showing that useful reader competence survives independently of the degraded adapted acquisition behavior. We call this a competence–policy transfer gap: adaptation improves evidence-conditioned competence, but that competence is not reliably translated into the decisions required to acquire useful evidence. Functional decoupling does not make inherited acquisition optimal. K2D-ALIGN (Knowledge-to-Decision Alignment) addresses the residual decision problem: when should retained reader competence justify a beneficial deviation from inherited evidence seeking? Matched continuation identifies Base-relative opportunities for learning root proposals; deployment-time admission authorizes replacement or executes exact Base fallback. Admission compares candidate and Base terminal shadow continuations before execution, trading additional model-side inference for conservative replacement. On HotpotQA, K2D-ALIGN improves Interactive EM from 46.1% to 52.3%; under matched initialization, Base-relative supervision contributes a further 3.9 points over the zero-update proposal. On 872 held-out cases, the frozen system gains 2.3 EM points over Base+FE and exceeds an approximately compute-matched Base Best-of-2 control by 3.1 points. Across four evidence-seeking benchmarks, it improves macro-average EM/F1 from 30.7/41.5 to 32.6/43.4. These results connect functional diagnosis to localized acquisition repair while preserving the reader competence that already transferred.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.