ClarifyAct: Learning to Clarify When Needed and Act When Ready
Abstract
Language models must not only know how to solve tasks, but also determine whether a request provides sufficient information to proceed. Missing premises, unclear intent boundaries, or erroneous constraints can cause failures on otherwise solvable tasks, whereas indiscriminate clarification creates unnecessary interaction costs. We define the performance difference between complete and defective requests as the suppression gap. Controlled experiments across multiple models, using complete, defective, and information-restored requests, empirically demonstrate this gap and recovery after information restoration. Motivated by this recoverability, we propose View-Specific Group-Relative Optimization (V-SGRO), which learns protective decisions from self-generated rollouts under task-verifier feedback, without external demonstrations. V-SGRO distinguishes verification status from relative quality, combining certification, preference, and residual quality signals into a composite advantage for unified policy updates. A fixed-candidate analysis establishes conditions for joint directional progress across these objectives. Experiments show that the full configuration improves the accuracy of decisions between protection and execution while reducing premature execution. Ablations and behavioral attribution analyses further reveal distinct effects of the learning signals on challenging requests.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.