More Compute Is Not More Evidence: Mitigating Model-Specific Error Bias at Test Time
Abstract
Test-time scaling methods implicitly assume that additional computation from the same model provides independent evidence. We show that this assumption fails within a single reasoning model: errors are correlated across repeated samples, across reasoning depth, and in the model’s own self-evaluation, so more compute from the same model is not more evidence. We introduce **DIA (Disagreement-triggered Independent Arbitration)**, which treats disagreement between an intermediate-stage and the final-stage prediction as a per-question signal of reasoning instability and resolves it by asking a small panel of heterogeneous models to solve the question independently, without seeing either candidate, and aggregating their answers directly. Across seven benchmarks and two base models, DIA outperforms full chain-of-thought in **12 of 14 settings**, including a **7.00% gain on MuSiQue**, while activating on only ** 17% of examples** and using **up to 47.2% fewer tokens** than a multi-agent baseline. Ablations show that both independent solving and solver heterogeneity are necessary for these gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.