acceptodds
Under review as a conference paper at ICLR 2027

Already Solved? When to Stop and What to Return in Iterative Reasoners

Abstract

An iterative reasoner can find a correct answer, keep computing, and ultimately return the wrong one. We study two coupled decisions: when to stop refining a state, and which generated answer to preserve. Our approach starts from a geometric observation: a hidden state can keep moving inside a region that decodes to the same answer. Margin-Motion Inclusion (MMI) compares the protection provided by the minimum logit margin with estimated remaining motion above an estimated noise floor, including a calibrated margin reserve against wrong early returns. We derive a decoding-ball guarantee without noise and conditions for recognizing the attractor's answer under persistent noise. Calibration then controls wrong early returns under exchangeability. The policy accepts and returns the first candidate that passes the test, potentially saving both iterations and restarts; otherwise, it returns the largest-margin candidate encountered. An oracle analysis separates missed solutions from delays in returning them. On recorded Sudoku and Maze trajectories from frozen Equilibrium Reasoners, fast MMI with maximum-margin fallback uses and fewer outer updates on average than terminal voting over 32 restarts of 64 updates. It also reduces errors by 20% and 75%, respectively. The gains come from recognizing and retaining answers, without changing the reasoner's learned dynamics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.