acceptodds
Under review as a conference paper at ICLR 2027

Epistemic Uncertainty for Test-Time Discovery

Abstract

Verifier-guided test-time search with large language models seeks one exceptional program within a small budget. TTT-DISCOVER trains the generating policy at test time with an entropic objective that emphasises the best sampled candidates, yet it can only credit programs the policy actually samples. We ask whether disagreement among policy adapters is a useful additional exploration signal. Uncertainty-Guided Test-Time Training (UG-TTT) maintains a small ensemble of low-rank adapters over a frozen base model, measures per-token disagreement as the mutual information between the next token and the ensemble member (an epistemic-uncertainty proxy, not a calibrated posterior estimate), and adds a rollout-level aggregate of it to the advantage as an exploration bonus. At a matched budget of 384 rollouts on six TTT-DISCOVER tasks, UG-TTT attains the higher maximum reward in five of the six seed-42 comparisons and a 0.03% lower one on the Erdos problem. Circle-packing differences are positive at all ˝ three seeds (with differing comparator configurations), outcomes on an autocorrelation inequality and a heuristic-contest task are mixed across seeds, and the single-cell denoising verifier’s error falls by 8.7% on Qwen3-8B and by 2.9% on a Qwen3-14B pair. In a one-run-per-arm component study on circle packing, the mutual-information bonus attains the highest maximum reward, and removing the adapter regulariser leaves disagreement nearly unchanged, indicating that diversified initialisation supplies the signal over six optimiser steps. Diversity depends on the window and the measure: at the end of training every seed-42 mathematical baseline has concentrated on one solution family while UG-TTT keeps at least three, and a rule-free structural measure favours UG-TTT in 9 of 12 run pairs over the late epochs and in 7 of 12 over the whole run. Ensemble disagreement is thus a practical exploration signal for budgeted test-time search; whether it beats simpler uncertainty bonuses, and whether diversity mediates its gains, remain open.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.