Do Agents Do What Users Ask? Measuring Intent-State Consistency in Tool-Using Language Agents
Abstract
Language agents increasingly act on behalf of users by invoking tools that send messages, book services, and modify persistent environment state. Existing evaluations measure task completion, attack success, policy compliance, or final-state correctness, but do not directly test whether the effects produced throughout execution remain within the user's authorization. We introduce contract-conditioned intent-state consistency, which evaluates agent behavior at three complementary levels: tool-call authorization, final-state consistency, and trajectory consistency. Starting from the same initial state, we replay reference and agent executions, associate tool calls with their realized state transitions, and identify missing required effects, persistent unauthorized effects, and transient unauthorized effects. Across 1,600 replayed AgentDojo trajectories from four models, these signals provide behavioral evidence complementary to task utility and attack-success labels. On a stratified human-review sample of 300 trajectories, full-trajectory evaluation achieves 0.814 precision, 0.809 recall, and 0.811 F1 for identifying executions judged inconsistent with user authorization, compared with 0.788 F1 for final-state-only evaluation. Trajectory replay further exposes intermediate effects that disappear before the final state. These results support intent-state consistency as a complementary diagnostic axis for evaluating tool-using agents under a supplied intent contract.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.