RefreshTrust: Provenance-Gated Decisions for Temporal Reward-Model Updating
Abstract
Temporal reward-model updating is an action-selection problem: channel-aware refresh can exploit a reliable collection trace, while unsupported channels can misallocate a fixed labeling budget as annotators, rubrics, tasks, and routing policies change. We introduce RefreshTrust, an auditable benchmark that selects among Full Trace-Aware Modular Refresh (TAMR), Observed-only TAMR, and Partial Refresh from the current provenance state. RefreshTrust matches backbone, 10% fresh-label budget, chronological split, tuning allowance, and held-out evaluation; stratifies channels as directly observed, audited proxy, or reconstructed; and freezes regime assignments before outcomes are examined. TAMR's additive channel decomposition and an architecturally distinct provenance-gated mixture of adapters expose complementary instruments for testing the same action map. On direct-metadata SHP-2 and T\"ulu-3-PREF anchors, Full TAMR gains accuracy points over Metadata-aware Refresh; across seven in-domain corpora the gain is points (), with over parameter-matched monolithic and continual-preference comparators. The assigned action beats Metadata-aware Refresh in 34/36 frozen cells under the TAMR instrument, 31/36 matched cells under PGMA, and 22/26 unperturbed natural cells; the probes' outcome-optimal actions agree in 33/36 cells. These results establish provenance quality as an operational control variable that determines how much temporal-refresh structure to use.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.