Learning Whether to Learn: Prospective Counterfactual Test-Time Learning for Language Models
Abstract
Test-time learning (TTL) promises models that improve during deployment, yet the dominant design pattern is myopic: an experience produces an update, and the update is judged by whether it fits that same experience. This creates a structural failure mode under unlabeled or weakly supervised streams—updates can reduce the current loss while degrading future decisions, amplifying pseudo-label noise, interference, and prompt-injection effects. We propose Foresight-TTL, a prospective formulation in which a candidate update is valued by its counterfactual effect on future deployment utility. During meta-training, matched rollout pairs estimate the counterfactual utility difference between applying and skipping the same update. A compact value model learns this future utility from the current experience, update geometry, uncertainty, and interference features. At deployment, a lower-confidence-bound gate decides whether to learn, a plasticity router decides where to write, and a reversible fast-weight transaction decides whether the write should survive. We further introduce multi-timescale persistence, promoting only repeatedly useful updates from episodic to persistent state. We analyze how calibrated value intervals can bound the probability of accepting harmful updates under a simultaneous calibration assumption. We evaluate the approach primarily on DeepSeek-R1-Distill-Qwen-7B across AIME 2024/2025, MATH-500, GPQA-Diamond, BIG-Bench Hard, and ARC-AGI-2, and additionally test cross-backbone deployment behavior on Qwen3-8B and Llama-3.1-8B-Instruct, with non-IID streams, corruption stress tests, and retention probes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.