RESOLUTION LIMITS OF A CONTINUAL-LEARNING BENCHMARK FOR LANGUAGE MODELS
Abstract
Small changes in continual learning are difficult to reliably distinguish from score variation due to task order, retraining seed and evaluation sampling. We quantify comparison reliability in a five-task classification benchmark with three language model backbones, trained under the same adaptation procedure. Across changes within the order, pooled rankings remain consistent, but magnitude and direction changes. A frozen fresh-seed run continuation verifies development agrees transfers or not. At five seeds per disjoint group, 13 out of 16 continue to meet a criterion of 95% sign agreement across seeds, splitting this holdout into replay vs non-replay comparisons gives us 9/9 and 4/7 agreeing comparisons respectively. This endpoint concerns agreement within each cohort, not preservation of the old winner or a replication probability. In a separate experiment-60 trajectories study, we find examples of close non-replay comparisons that change within fixed execution environments. Mixing hosts is thus not the only source of variation. Conditioned on fixed trained models, jointly resampling evaluation datasets across all training and testing stages leads to a median evaluation-sampling standard deviation of 0.0119, versus 0.0158 when resampling at each stage independently. We provide as part of the audit an executable that will reproduce these paired effects, uncertainty, declarations, abstentions and the observed cost of selective reporting. Reported near-superiority on new runs cannot be certified without assessing each method’s difference and each reporting decision. These results hold for this research, but do not establish universal seed requirement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.