acceptodds
Under review as a conference paper at ICLR 2027

BEYOND ENDPOINT SCORES: WHEN CONTINUAL KNOWLEDGE-UPDATING RANKINGS REVERSE

Abstract

Continual knowledge-updating methods are usually compared at a single final checkpoint. On a retrospective 24-update factual stream we show that this reporting choice, the initialization of the replay baseline, and the query form can each reverse the outcome of one and the same comparison. The testbed pairs a fast/slow LoRA hierarchy—a slow adapter re-fitted every K updates plus a small monthly adapter—against cumulative LoRA replay on Llama-3.2-1B and Qwen2.5-1.5B. Baseline: carrying replay's adapter weights across updates raises held-out paraphrase retention by 12.8 points at rank 8 and turns the hierarchy's +7.3-point full-history lead into a −5.4-point deficit, while lowering unchanged-fact accuracy by 12–19 points. Window: because the final update is a consolidation state for every tested K, the endpoint favors the hierarchy by 9–11 points even where warm replay leads by 5–18 points on average; the reversal appears in 16 of 18 seed-level paraphrase comparisons. Query form: training-form queries sit near ceiling under warm replay and cannot discriminate. An exact decomposition of reporting error into within-cycle phase and between-cycle drift explains why averaging the last cycle still selects the wrong winner in four of 18 cases. A retrospective drift check passes in seven of 12 aggregates and withholds the one wrong last-cycle sign, and spreading the same number of checkpoints over the stream removes every sign error on the warm panel. We distill these findings into a reporting protocol for continual-updater comparisons.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.