acceptodds
Under review as a conference paper at ICLR 2027

A Permutation Audit for Cell-Level Foundation-Model Transferability Claims

Abstract

Transferability scores such as LogME are sometimes read as cell-level warnings: a low score for a (foundation model, task) cell is taken to predict that this particular cell will go wrong. The usual evidence is a correlation between the score and cell-level error across a grid of models and tasks. Such a correlation can be produced without any cell-level information, by shared task difficulty, by shared model strength, or by the noise in the error estimate itself. We give an audit for such claims: gates that check that the score is defined on the cells and that the outcome has cell-level variance not dominated by seed noise, then three within-task permutation nulls, each aimed at one artefact, all of which must reject in the claimed direction. The noise null adjusts for each cell's measured seed variation instead of dividing by it; we show that a natural ratio standardisation rejects genuine signal when noise grows with difficulty (pass rate ), whereas the adjusted null does not (1.00). On a synthetic grid the audit passes 0.6% of no-signal runs and none of five leakage regimes, and detects genuine within-task correlations of 0.4 in every run. Applied to LogME, the H-score and N-LEEP on seven foundation-model pools, 11 of 84 warning-oriented audits pass, 10 of them on two medical-imaging pools where two of the three scores pass on the same outcomes; the audit also identifies where it cannot decide: outcomes dominated by seed noise, targets without seed replicates, and grids too small to have power. The implementation returns a structured record that can accompany a new claim.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.