acceptodds
Under review as a conference paper at ICLR 2027

ODAR: Instance–Compute Alignment for Adaptive Test-Time Reasoning

Abstract

Test-time scaling is commonly framed as allocating more reasoning to more difficult inputs, yet static difficulty need not track the realized benefit of additional computation. We present ODAR, an adaptive reasoning system whose deployed router learns from coarse complexity and Fast-Agent failure signals, while all-path counterfactual outcomes are reserved for evaluation. In an all-path audit, this inexpensive routing proxy predicts positive compute benefit with 0.77 AUROC, compared with 0.58 for input length, 0.68 for predictive uncertainty, and 0.694 for a PRM score. Holding the Simple/Medium/Hard proportions and average calls fixed, within-benchmark permutation lowers audit performance by 4.29 points, isolating instance–compute alignment; frozen-router transfer retains 74.4–87.8% of source-pair gains. Across 22 public benchmarks plus one custom proof stress test, ODAR-API exceeds Self-Consistency on all 23 displayed task-specific metrics while averaging 2.55 model calls. Token-level FEP selection is evaluated in the open-weight configuration, where next-token likelihoods are observable. Together, these results establish realized compute benefit as an informative evaluation target for adaptive inference and show that inexpensive supervision can identify instances that benefit from additional reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.