acceptodds
Under review as a conference paper at ICLR 2027

No Single Point Suffices: Reporting Regularization Landscapes in Frozen-Representation Evaluation

Abstract

Frozen-representation evaluation compares pretrained representations through probe performance, but the measured advantage depends on the evaluation protocol. We study signed comparison gaps with ridge probes on three music benchmarks and an exploratory NLP/vision extension, distinguishing shared-penalty paths from independent validation-based selection. In the primary MTG-Jamendo configuration the sampled ROC-AUC gap has range 0.04432 and endpoint contrast 0.04156 (nominal 95% artist-cluster-bootstrap CI [+0.00765, +0.07980], 10,000 resamples), remaining positive at all ten sampled coefficients; simultaneous bands admit a constant gap curve. Both inherit the declared grid's extent: trimming the upper bound to the two arms' sampled peak reduces the endpoint contrast to +0.0060 and the range to 0.00877. Within the gap is not distinguishable from a constant gap under the paired cluster bootstrap (); the movement the bootstrap resolves is the post-peak rise from to . Restricting each arm to validation-near-optimal coefficients reveals a different reporting limitation: at , two of 22 cross-domain comparisons — each with a singleton common-coefficient set of zero range — admit independently acceptable test gaps of both signs, while the other twenty remain sign-consistent; matching effective complexity leaves four of 22 comparisons sign-changing. These exploratory envelopes are descriptive, not uncertainty-adjusted ranking claims. Different evaluation questions are therefore not interchangeable: a narrow shared-coefficient summary can omit variation under independent choices.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.