When Do State-Based Diagnostics Transfer in Soft-Prompt Tuning?
Abstract
Understanding how adaptation changes language-model behavior is important for reliable evaluation and deployment. In soft-prompt tuning, we ask whether an internal diagnostic's advantage over answer loss transfers across prompt groups and output rules, and whether its ranking gains justify measurement costs. We propose a control-response evaluation framework to address these questions in frozen language models. First, in two Ministral–ECQA prompt pairs, removing continuation reverses the ranking of state and loss diagnostics under strict parsing, while preserving diagnostic scores and initially selected answers. This shows that detecting parseability changes can drive an apparent diagnostic advantage. Second, crossed prompt–question comparisons in Ministral–ECQA show larger changes in relative diagnostic performance across prompt groups than across question sets. The group contrast includes pair selection. Predictions fixed before confirmation reproduce this shift on fresh questions for the same prompts and improve on predicting no group difference. Pair-specific gains over a common-mean forecast remain uncertain, as do gains on unseen GSM8K prompts. Third, at the measured ECQA reference budgets, ranking gains do not offset checks lost to scoring, and direct checks outperform screening with the three original diagnostics on average. The framework provides a reproducible procedure for interpreting diagnostic advantages, testing their transfer, and deciding whether internal measurement is worth its evaluation cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.