One Model, Many Minds: Strategy-Diversified Reasoning Is a Cross-Model Ensemble in Disguise
Abstract
The dominant recipe for test-time scaling is to sample more: draw many reasoning chains from one model and aggregate them. Yet chains from the same model under the same prompt share the same failure modes, so their errors are correlated and accuracy saturates well before the compute budget does. Cross-model ensembles break this correlation, since different architectures fail differently, but at the price of running several large models. We ask whether that diversity is genuinely architectural, or merely structural and reproducible within a single model. By forcing one model through deliberately distinct reasoning strategies (algebraic, geometric, backward, inductive, analogical), we decorrelate its errors most of the way toward switching models. Across four diversity sources, the mean pairwise error correlation follows a clean hierarchy: temperature () and prompt rephrasing () keep errors tightly coupled, whereas strategy diversification drops them to , a third of the way to a five-model ensemble (), without lowering per-attempt accuracy, and so attains the highest answer coverage of any source. At matched compute, the resulting single-model ensemble, One Model, Many Minds (OMMM), matches or beats the compute-matched baselines and rivals far costlier methods such as multi-agent debate. The decorrelation also buys calibration: because the strategies fail independently, the rate at which they agree is a free, well-calibrated confidence signal, and answering only when the strategies agree raises accuracy on the answered AIME problems from 37% to 100% (a selective-prediction gain, at reduced coverage). What decorrelates a model's errors is not how many times it answers but how differently it reasons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.