acceptodds
Under review as a conference paper at ICLR 2027

Which Questions Are Worth Thinking About? Cross-Developer Transfer of Thinking Value

Abstract

Which questions are worth the additional cost of thinking? We test whether question-level accuracy gains from configured thinking modes transfer across model developers. We define thinking value as the paired accuracy change between configured thinking and non-thinking modes and evaluate 18 models from seven developers on 300 MMLU-Pro STEM questions and 198 GPQA-Diamond questions under three prompting strategies. Gains estimated from other developers predict a held-out developer's gains after IRT-style adjustment for measured difficulty and discrimination, with pooled correlations of 0.264 and 0.273. Transfer persists with an independent difficulty panel and under request-setting and scoring checks. In retrospective allocation on the evaluated questions, ranking transferred gains per predicted additional token recovers 91.8% and 79.1% of the always-thinking accuracy gain at half of the always-thinking additional-token budget, compared with 83.5% and 74.4% for a difficulty-based gain predictor per token. The advantage is statistically resolved on MMLU-Pro but not GPQA-Diamond. An ability-shift simulation produces comparable transfer, leaving the source of this predictability unresolved. The tested text-based predictors generalize weakly to unseen questions. These findings characterize cross-developer transfer of thinking value on previously evaluated questions and motivate learning to predict this value for adaptive computation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.