acceptodds
Under review as a conference paper at ICLR 2027

Goodhart-Aware Test-Time Reasoning

Abstract

Recent work has shown that injecting learnable latent tokens into frozen large language models (LLMs) and optimizing them at test time by maximizing a confidence-based proxy reward improves reasoning performance. However, we show that optimizing such rewards without bound can lead to Goodhart overoptimization, where accuracy peaks at a problem-dependent threshold and degrades beyond it. We propose GoATT-R (GoodhartAware Test-Time Reasoning), which treats the frozen LLM as an uncertain system and steers its output using an Ensemble Kalman Filter (EnKF) that tracks a prescribed proxy target reward rather than maximizing the proxy indefinitely. To set this target adaptively, GoATT-R employs a lightweight online policy network, trained via SGD during evaluation, that predicts a suitable optimization target from prompt features and initial ensemble statistics using majority-vote agreement as a label-free reward signal. GoATT-R consistently outperforms state of the art baselines across multiple language models and reasoning benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.