Say It Differently, Act Differently? Metamorphic Testing of Operational Behavior Consistency in AI Agents
Abstract
LLM-powered AI agents increasingly perform real-world tasks through multi-step operations, making operational consistency an important aspect of reliability. Yet existing evaluations typically consider only a single expression of a user request, overlooking the many ways in which the same intent may be expressed. This raises a key question: do intent-equivalent requests lead to operationally consistent agent behavior? To investigate this question, we propose AgentMT, an MT-based framework for evaluating black-box AI agents under intent-preserving prompt variations. AgentMT transforms prompts along tone, language, and formulation, executes them under the same agent setting, and represents observable executions as operation graphs to compare permission decisions, information disclosure, and execution trajectories. Across three datasets and commercial AI agents, follow-up prompt variations induce operational inconsistency in all 27 model–dataset–transformation combinations, reducing mean operation-graph similarity by up to ; significant task-level consequence inconsistency is observed in 22 of 27 model–transformation combinations. We further assess generalization using GPT-6 Astra, multi-turn clarification, Claude Code, and Codex CLI. Our findings show that preserving user intent does not guarantee consistent agent operation, motivating systematic evaluation of operational robustness across diverse user expressions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.