Semantic Replay Has a Coordinate System: Claim-Level Evaluation for Delayed-Feedback RL
Abstract
A replay buffer can preserve every transition yet change the meaning of its stored state when an encoder update moves the representation coordinates. Delayed semantic-history RL compounds this problem: information-time, censoring, system attribution, and resource conclusions can require different evidence even when they share one return. TRACE (Temporal Replay Admissibility and Claim Evaluation) makes the evidence unit explicit—a policy row or a registered contrast—and maps deterministic provenance failures, scope projections, and residual assumptions to the claims each unit can carry. The operator has absorbing hard failures, commutative claim projections, and explicit empty-set semantics. On WebShop, checkpointed replay yields 0.340 success when old coordinates are reused, 0.398 with version-filtered sampling, and 0.402 with full re-encoding; filtering recovers 93.5% of the point-estimate gap at 0.38 times the re-encoding cost. Under the same 500K-step protocol, the evaluated LLaMA-2-7B and pretrained RoBERTa-base systems reach 0.398 and 0.386 \pm 0.018 WebShop success, above the Causal Transformer's 0.358 and GRU's 0.312 point estimates; on ALFWorld they reach 0.375 \pm 0.020 and 0.364 \pm 0.019, with a teacher–Delay-Survival difference of 0.043 (95% CI [0.017, 0.069]). Post-decision positive controls score still higher on both language workloads but are excluded by field provenance, while RecSim's 0.041–0.046 conversion range remains linked to its censoring model. The result is a claim-specific account of semantic replay: a practically large replay-strategy effect, replicated system-level performance ordering across two language-mediated workloads, and conclusions whose statistical and observational conditions remain visible.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.