acceptodds
Under review as a conference paper at ICLR 2027

Reliability-Guided Objective Allocation for Test-Time Reasoning

Abstract

Test-time adaptation can improve large language model reasoning without additional supervision, yet existing approaches typically apply the same optimization objective to every input. We argue that this uniform treatment ignores a fundamental asymmetry in test-time learning: reliable solutions should be consolidated, whereas uncertain ones should be explored. We introduce DiSCTT, which uses agreement among independently sampled reasoning trajectories as a label-free estimate of instance-level reliability and dynamically allocates the adaptation objective accordingly. High-consensus inputs are routed to supervised fine-tuning using modal-answer-supported pseudo-targets, while low-consensus inputs are routed to reinforcement learning with a modal-consistency-gated, population-aware reward. The routing partition is periodically recomputed as the model adapts, yielding a self-evolving curriculum between consolidation and exploration. Across diverse reasoning benchmarks and models ranging from 0.5B to 32B parameters, DiSCTT consistently improves over uniform reinforcement-learning adaptation, including gains of 9.6 and 8.7 percentage points on the three-task average for 20B and 32B models, respectively. Controlled comparisons against random and static routing, sequential SFT+RL, a label-free comparator based on SRFT (LF-SRFT), and an external-verifier variant support consensus-conditioned objective allocation under the evaluated configurations. DiSCTT also substantially reduces adaptation compute relative to full-dataset reinforcement learning and extends to open-ended generation through semantic consensus. These results suggest that test-time adaptation should treat the choice of optimization objective itself as an instance-level decision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.