acceptodds
Under review as a conference paper at ICLR 2027

A Protocol for Auditing Replay-to-Online Transport in Recursive Self-Improvement

Abstract

Historical replay can identify a promising strategy without showing how that strategy will behave when new tasks arrive. We present a registration and reporting protocol for this replay-to-online boundary. The protocol freezes the candidate set and replay winner before revealing a separately registered live block, limits each tested strategy to the input allowed for the current task, and compares the winner with the incumbent on the same tasks and registered repeats. It reports replay gain ΔR, paired live gain ΔL, and their descriptive difference B = ΔR - ΔL. A two-world construction shows why the same replay record can be consistent with opposite live gains when replay and live outcomes are otherwise unrelated. The checker rejects duplicate task identities, incomplete replay coverage, and declared protocol violations; unsupported or missing observations remain UNKNOWN or PARTIAL. We include synthetic checks, a deterministic fixed-policy system check, and a finite model-assisted pilot. These artifacts test the contract and illustrate its reporting boundary; they do not establish predictive validity, population generalization, statistical significance, or online performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.