Under review as a conference paper at ICLR 2027
Separating Answer Selection from Computation in Adaptive Language-Model Policies
Abstract
A confidence-guided policy can improve answered accuracy by choosing which answers to return as well as by allocating additional reasoning. We test a fixed controller through interventions that preserve cached branch outcomes, exact query counts and answer coverage. Across two open-weight backbones on MMLU and ARC Challenge, the 25%-query, 80%-coverage controller exceeds state scrambles by 7.50 macro accuracy points. Query-only and abstain-only controls yield gains of 1.55 and 6.94 points at their respective budgets. The results establish that item-aligned state matters for this controller under fixed budgets.
open until 14 Dec 2026
est. 32% chance this paper gets accepted at ICLR 2027.
Reject 68%Accept 32%
What do you think this paper will get?
All positions stay anonymous.
Related papers
Loading the map…
Discussion (0)
Sign in to comment.