Adapt Only When It Pays: Selective, Verifiable Test-Time Self-Training for LLM Reasoning
Abstract
Test-time optimization adapts LLM parameters during inference and can correct misconceptions that test-time scaling cannot; recent query-conditioned approaches remove the dependence on external data by synthesizing structurally related problem–solution pairs from the input query itself and fine-tuning on them before answering. This recipe is effective but uniform in two respects: adaptation is **unconditional** — every query pays the full synthesis-and-update cost, including the majority the base model already answers correctly — and every synthesized pair is weighted **equally as supervision**, although it originates from the same model whose competence is in question, a reliability gap prior work identifies as a limitation but does not instrument. We propose **SelVer-TTT**, which makes test-time adaptation selective and its supervision verifiable: a cheap pre-adaptation probe gates the expensive update to queries that are both uncertain under the base model and resolvable by adaptation; for gated queries we bias synthesis toward variants with automatically checkable structure and weight each pair by whether its solution survives verification, so unverified samples never act as positive supervision; and following the persistent test-time optimization direction outlined as future work by prior query-conditioned self-training, we accumulate LoRA increments into a **bounded, decaying residual slot** under an explicit norm budget, reporting the resulting drift directly. Across 7 mathematical reasoning benchmarks and GPQA-Diamond, SelVer-TTT matches or exceeds unconditional query-conditioned self-training while adapting on only 38% of queries and reducing average wall-clock inference overhead by 2.4×, with verification-weighted supervision contributing 1.9 points over equally-reliable weighting and widening on the hardest split; the bounded residual memory retains 71% of per-query adaptation benefit across a 500-query stream while keeping degradation on a held-out general benchmark below 1.2 points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.