acceptodds
Under review as a conference paper at ICLR 2027

More Compute Is Not More Evidence: Mitigating Model-Specific Error Bias at Test Time

Abstract

Test-time scaling methods implicitly assume that additional computation from the same model provides independent evidence. We show that this assumption fails within a single reasoning model: errors are correlated across repeated samples, across reasoning depth, and in the model’s own self-evaluation, so more compute from the same model is not more evidence. We introduce **DIA (Disagreement-triggered Independent Arbitration)**, which treats disagreement between an intermediate-stage and the final-stage prediction as a per-question signal of reasoning instability and resolves it by asking a small panel of heterogeneous models to solve the question independently, without seeing either candidate, and aggregating their answers directly. Across seven benchmarks and two base models, DIA outperforms full chain-of-thought in **12 of 14 settings**, including a **7.00% gain on MuSiQue**, while activating on only ** 17% of examples** and using **up to 47.2% fewer tokens** than a multi-agent baseline. Ablations show that both independent solving and solver heterogeneity are necessary for these gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.