Correct Numbers, Wrong Constructs: Acquisition-Aware Evaluation Across Model Scales
Abstract
We built a program to learn, at a small scale, which adaptations make a model forget, and to test whether that knowledge survives an order of magnitude of scale. It returned a striking headline: roughly 98% of the relationships it selected failed to transfer. This paper is the audit that stopped us reporting it, and what the audit found was not one mistake but five instances of one. Each has the same shape: a quantity was computed correctly, and a different quantity was inferred from it. A count of records read as a count of independent observations, where 186 "pairs" were 62 directed pairs recorded three times, byte-identical across all 558 matched groups. A failure rate read as evidence that selection failed, at 98.7% against 98.8% for choosing at random on the same population. Deterioration read as forgetting, after an update that trained repeatedly on one 32-sequence batch and left held-out loss worse in 60 checks across three scales. A shared recipe's collapse read as a property of scale, when a per-scale recipe acquires in every confirmation cell and a width-scaled learning rate alone moves 1.4B from +0.043 to +0.239. And a preregistered median read as agreement between task families, where with three families the median is the middle family's value, certifying a match that held at a 0.0019 nat spread in one family while another disagreed by 0.5079. None required an unusual mistake, and preregistration caught none of them because one of them is the preregistration. A sixth arrived while we built the corrected experiment, this one caught prospectively: the per-scale recipe our own median rule certified for 160M acquires in just 32 of 62 pairs at that scale. We give an acquisition-aware protocol and a checklist with one item per substitution. Across 18 published papers a screen finds an explicit held-out new-task metric in 11 and does not resolve the other 7, so the defect is not universal and the screen does not establish how far it reaches. Asked at last under a preregistration written before any fit existed, the original question does not resolve: our candidate reaches a held-out R^2 of -0.004 and fails its margin against the best baseline, which is the one comparator with a positive point estimate, +0.023. On 15 held-out pairs we cannot separate absence of signal from absence of power, and we report both.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.