Before Commit: Purpose-Aware Rehearsal for Tool-Using Agents
Abstract
Tool-using agents can produce actions that are syntactically valid, policy-compliant, and executable, yet persistently alter external state without advancing the user's current purpose. Existing safeguards can improve how agents plan, reason, and recover, but they do not necessarily verify the final reversible decision before an external write is committed. We introduce Purpose-Aware Rehearsal (PA-Rehearsal), a training-free, candidate-bound interface that intervenes at this pre-commit boundary. For each proposed write, PA-Rehearsal binds the exact candidate to the latest user purpose and observable evidence, and uses an LLM to rehearse the action's likely effect and the goal it would serve. A deterministic controller retains final release authority by enforcing mandatory evidence, resolving conflicts between rules and model judgments, and verifying candidate identity. The interface leaves read-only exploration unchanged, rechecks every revised candidate, and can be layered onto execution-time or post-execution assistance. In matched three-run evaluations on -bench Airline and Retail, PA-Rehearsal attains the highest pass@1 point estimates among the compared methods. Notably, it outperforms a prompt-only boundary checklist by 20.0 percentage points on Airline. On WebShop, a benchmark for grounded web interaction, adding pre-purchase rehearsal to a fixed trajectory-regulation scaffold raises pass@1 from 44.5% to 50.0%. Together, these results show that purpose-aware checking at commitment boundaries can improve reliable task completion, positioning pre-commit rehearsal as a distinct, composable control primitive for tool-using agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.