Test-time RL alignment exposes task familiarity artifacts in LLM benchmarks
Abstract
Direct evaluation of LLMs on benchmarks can be misleading because different models may have acquired different degrees of familiarity with the task prior to evaluation. The train-before-test approach controls for task familiarity by giving each model task-relevant training before evaluation, originally through supervised fine-tuning(SFT). However, suitable training data is often hard to come by, and evaluation results vary with the data chosen. In this paper, we propose a two-stage test-time reinforcement learning (RL) alignment method for train-before-test. First, RL with a single sample provides a first alignment of the model to the task format, and second, test-time RL with majority-voting reward aligns the model to the benchmark distribution. Our test-time RL alignment method aligns similarly well as the SFT-based train-before-test, but without requiring a task-specific training set. On a domain-specific benchmark without training data, we show that direct evaluation underestimates base models, which perform substantially better once aligned, yielding a comparison of models under the same task preparation. Moreover, for reasoning tasks, base models reach performance comparable to many of their fine-tuned variants after the same label-free test-time alignment, while the fine-tuned variants themselves exhibit only marginal changes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.