Which Questions Are Worth Thinking About? Cross-Developer Transfer of Thinking Value
Abstract
Which questions are worth the additional cost of thinking? We test whether question-level accuracy gains from configured thinking modes transfer across model developers. We define thinking value as the paired accuracy change between configured thinking and non-thinking modes and evaluate 18 models from seven developers on 300 MMLU-Pro STEM questions and 198 GPQA-Diamond questions under three prompting strategies. Gains estimated from other developers predict a held-out developer's gains after IRT-style adjustment for measured difficulty and discrimination, with pooled correlations of 0.264 and 0.273. Transfer persists with an independent difficulty panel and under request-setting and scoring checks. In retrospective allocation on the evaluated questions, ranking transferred gains per predicted additional token recovers 91.8% and 79.1% of the always-thinking accuracy gain at half of the always-thinking additional-token budget, compared with 83.5% and 74.4% for a difficulty-based gain predictor per token. The advantage is statistically resolved on MMLU-Pro but not GPQA-Diamond. An ability-shift simulation produces comparable transfer, leaving the source of this predictability unresolved. The tested text-based predictors generalize weakly to unseen questions. These findings characterize cross-developer transfer of thinking value on previously evaluated questions and motivate learning to predict this value for adaptive computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.