Test-Time Revision: Latent Reasoning with Execution in the Loop
Abstract
Extra test-time computation improves predictions. However, updating hidden states during inference may improve intermediate scores yet still produce incorrect or invalid outputs after full task execution. Reranking can select a better output, but selection alone does not ensure that subsequent reasoning builds on the corresponding internal state. We introduce *Test-Time Revision (TTR)*, an approach to latent reasoning guided by task execution. TTR evaluates bounded changes to the current state through full task execution and jointly commits the selected output and its corresponding internal state. This preserves consistency between the selected output and the state used for subsequent reasoning, while allowing execution feedback to guide further revision. The same mechanism applies to Sudoku token states, Maze spatial representations, and Qwen prompt embeddings. Each task retains its own adapter, executor, and evaluation rules, while the original backbones remain frozen. TTR achieves strong performance on held-out Sudoku, Maze, and MATH500 data, solving of Sudoku boards correctly. Matched Maze controls show that reusing the selected state improves subsequent execution, and the anonymous code is available at [https://anonymous.4open.science/r/4D-570D/](https://anonymous.4open.science/r/4D-570D/).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.