Scalar Recalibration Can Overstate Contextual Gains in Pairwise Prediction
Abstract
Does a contextual predictor add value if its advantage survives recalibration of a simpler model? In pairwise prediction, this conclusion can depend on whether the calibrator observes which competitor occupies a known side. An ordinary Platt intercept in blue-first or white-first coordinates is exactly a side-dependent intercept in randomly ordered coordinates. We use this identity to compare ordinary and side-aware Platt scaling while holding each model's forecast stream fixed. Across seven annual League of Legends splits and nine weekly Lichess splits, covering 70,670 and 1,340,775 evaluation matches, respectively, an online side term improves a calendar rating model's log loss after ordinary Platt scaling by 0.001144 and 0.000119. After both readouts receive side information, the differences reverse to +0.000139 and +0.000013, with 95% paired calendar-block intervals of [-0.000087,0.000325] and [0.000004,0.000024]. A four-arm League of Legends ablation also finds no proper-score benefit from the existing frozen rank-five interaction term; shared and independent validation selection yield identical forecasts. These results do not establish equivalence or rule out richer contextual models. They show that survival under scalar recalibration is insufficient evidence of contextual value beyond a simpler predictor with the same side information.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.