CloudIntent-Bench: Intent Multiplicity Challenges LLM Agents in Real-World Cloud Services
Abstract
Large language model (LLM) agents are increasingly deployed in interactive scenarios, such as cloud services, where conversations often involve multiple user requests, challenging the underlying LLMs to reason over evolving dialogue histories. Although existing analyses often attribute lower performance to longer dialogue histories, interaction turn counts alone provide limited insight into which aspects of a conversation drive this degradation. We therefore investigate **intent multiplicity**, defined as the number of distinct user intents identified in a dialogue history by intent annotations. It provides an intermediate characterization between raw conversational context and agent performance, enabling us to analyze agent performance in real-world applications at a more fundamental level. Furthermore, we introduce CloudIntent-Bench, a benchmark of decision tasks derived from authentic cloud service tickets, to systematically examine the LLM-agent performance in different intent multiplicity scenarios. Experimental evaluations show that performance declines as intent multiplicity increases within a transition region, and controlled studies that remove or provide intent information offer complementary evidence for attributing this decline to intent multiplicity. Our work identifies intent multiplicity as a more fundamental perspective for analyzing agent performance and provides a practical benchmark for developing reliable LLM-agent applications in the cloud services domain.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.