acceptodds
Under review as a conference paper at ICLR 2027

Separating Answer Selection from Computation in Adaptive Language-Model Policies

Abstract

A confidence-guided policy can improve answered accuracy by choosing which answers to return as well as by allocating additional reasoning. We test a fixed controller through interventions that preserve cached branch outcomes, exact query counts and answer coverage. Across two open-weight backbones on MMLU and ARC Challenge, the 25%-query, 80%-coverage controller exceeds state scrambles by 7.50 macro accuracy points. Query-only and abstain-only controls yield gains of 1.55 and 6.94 points at their respective budgets. The results establish that item-aligned state matters for this controller under fixed budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.