acceptodds
Under review as a conference paper at ICLR 2027

Say It Differently, Act Differently? Metamorphic Testing of Operational Behavior Consistency in AI Agents

Abstract

LLM-powered AI agents increasingly perform real-world tasks through multi-step operations, making operational consistency an important aspect of reliability. Yet existing evaluations typically consider only a single expression of a user request, overlooking the many ways in which the same intent may be expressed. This raises a key question: do intent-equivalent requests lead to operationally consistent agent behavior? To investigate this question, we propose AgentMT, an MT-based framework for evaluating black-box AI agents under intent-preserving prompt variations. AgentMT transforms prompts along tone, language, and formulation, executes them under the same agent setting, and represents observable executions as operation graphs to compare permission decisions, information disclosure, and execution trajectories. Across three datasets and commercial AI agents, follow-up prompt variations induce operational inconsistency in all 27 model–dataset–transformation combinations, reducing mean operation-graph similarity by up to ; significant task-level consequence inconsistency is observed in 22 of 27 model–transformation combinations. We further assess generalization using GPT-6 Astra, multi-turn clarification, Claude Code, and Codex CLI. Our findings show that preserving user intent does not guarantee consistent agent operation, motivating systematic evaluation of operational robustness across diverse user expressions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.