Does Prompt Sensitivity Improve Cross-Family Degradation Prediction?
Abstract
Models whose answers barely change when instructions are reworded might also hold their accuracy when the data change. Does prompt sensitivity predict how much accuracy a new model family will lose, beyond source accuracy? We test two source-side features, accuracy dispersion and label-free disagreement across ten wordings, on fifteen open-weight models, three classification tasks and seven natural shifts. Both lower held-out error under cross-validation over five development families. Under one frozen fit, disagreement raises error by 6% in a first held-out cohort and lowers it by 35% and 30% in a second and a third, of three and two families; dispersion has the same signs. No cohort reaches the registered level, and no test pools them. Disagreement's falls in the second and third cohorts exceed a within-source permutation reference, but in development the invalid-response rate alone predicts degradation nearly as well, and the first cohort's rise reverses when one training family is dropped. The evidence does not establish that the gain transfers to a new family, nor that sensitivity carries no information about shift.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.