acceptodds
Under review as a conference paper at ICLR 2027

Reference Accuracy and Predictive Value in Low-Precision Training

Abstract

Direct low-precision write-back can erase nonzero optimizer proposals. A high-precision reference can evaluate that event on a projected state, but aggregate agreement need not imply improvement over simpler controls. In a controlled grid, 55/72 cells have measured and predicted post-initialization crossings spanning ; 52/55 lie within 15%, while four cells disagree on crossing category. A reused-source NeoX audit passes temporal skill against a constant, yet an outcome-blind secondary comparison favors a historical template (RMSE 0.00360) over the corrected source (0.00438). A prospective Pythia-410M checkpoint-restart stage evaluates 15 holdout targets across three ages. Its arithmetic-margin selector meets study-specific absolute error/coverage requirements on every target but has higher error than always predicting unchanged on the same accepted events. Every unit-gain source correction also loses to its template. Weak residual alignment limits even a separate hindsight-optimal scalar gain to less than 3.4% RMSE improvement per observed comparison, below the fixed 10% requirement; this is a post-outcome geometric result. The complete original protocol already rejects the forecast, and finalization includes a disclosed analysis-binding repair. Matched rounding interventions separately show large fixed-run policy contrasts. These empirical results characterize this reference construction and its controls, rather than a general failure of reference forecasting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.