acceptodds
Under review as a conference paper at ICLR 2027

Can Language Models Learn from Failures Alone at Test Time?

Abstract

Test-time training (TTT), where language models (LMs) undergo parameter updates during inference, has emerged as a promising approach for adapting LMs to individual problems. Prior work on problem-specific TTT often operates in hill-climbing settings, where candidate solutions can be evaluated with a fine-grained scalar reward. This paper asks: how far can test-time training go when learning only from failures, without such fine-grained feedback? We adapt a range of existing optimization methods to this regime and propose Controlled Redistribution (CoRed), which treats failure-only TTT as controlled reshaping of the model's token-level predictive distribution. Across countdown games, math, and code tasks, CoRed improves both solvability and sample efficiency. On challenging math problems, CoRed enables Qwen3-4B-Instruct to solve more than as many problems as the frozen base model within 4,096 attempts, while substantially reducing the number of attempts required to reach the first correct solution on problems that the base model can eventually solve within 4,096 attempts. Our results demonstrate that substantial test-time adaptation is possible even when learning from failures alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.