acceptodds
Under review as a conference paper at ICLR 2027

Test-time Learning from Reasoning Failures

Abstract

When an attempt is rejected during deployment, large language models (LLMs) commonly retry through resampling or textual feedback for self-correction, while keeping parameters fixed. Yet a rejected attempt itself provides a negative learning signal: the model can update its parameters to reduce the attempt's likelihood, without adding it to the context. However, simply unlearning a rejected attempt before retrying does not guarantee better overall performance. To this end, we introduce Negative Space, a bilevel optimization framework that trains LLMs to adapt from rejected attempts at test time. Specifically, during training, the lower level simulates test-time adaptation from failure: it applies a single unlearning step to rejected responses between two successive attempts. The upper level learns to make this adaptation beneficial for overall performance through first-order policy optimization that maximizes rewards across initial and retry attempts. At test time, a rejected attempt triggers the same unlearning step; the model then retries the original query and discards the update. Across code generation and mathematical reasoning, Negative Space improves average accuracy over the strongest evaluated retry baseline by 4.8 points on two coding benchmarks and 3.7 points on five mathematics benchmarks. Further analysis shows that, although trained with fewer retries, it retains its advantage when given more at test time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.