acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Efficiency-Induced Premature Stopping in Knowledge-Intensive Agentic Reasoning

Abstract

Agentic reasoning is essential for knowledge-intensive tasks, where models must iteratively decide whether to continue reasoning, call a tool to retrieve external evidence, or produce an answer. However, these models often engage in lengthy deliberation before taking an action, substantially increasing token cost and latency without making meaningful progress. Existing test-time efficient methods mainly rely on efficiency-oriented early stopping, which may terminate reasoning before a necessary tool call and lead to severe performance degradation. To address this efficiency-induced premature termination, we propose OPERA, a lightweight test-time framework that uses only the reasoning model to reduce redundant deliberation while preserving productive reasoning. OPERA uses hierarchical action probing to assess whether the model is ready to take the next action, and identifies potential verbosity on-the-fly to estimate when to probe. It terminates reasoning only when probes consistently support the same action. Extensive experiments across knowledge-intensive benchmarks show that OPERA consistently reduces token usage while preserving or improving answer accuracy and more effectively mitigates premature termination than baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.