acceptodds
Under review as a conference paper at ICLR 2027

-Emotion-Bench: Exposing the Interaction Fragility of LLM Agents through Real-World User Emotions

Abstract

In long-horizon tool-calling tasks, LLM agents can fail not only because of tool-use errors, but also because of ineffective interaction with users. Yet existing benchmarks largely assume cooperative and behaviorally stable users, leaving the effect of emotional variation on agent reliability underexplored. We introduce -Emotion-Bench, a benchmark for evaluating tool-calling agents under emotion-conditioned user interactions across ten domains. To construct the benchmark at scale, we introduce Tracer, a scalable agentic data-generation pipeline that constructs policy-valid tool workflows, grounds them into executable tasks, incorporates user preferences and task-specific emotional context, and verifies task validity against the environment. The resulting tasks are mapped to a structured emotion space and paired with emotion-conditioned user simulators, enabling controlled evaluation while preserving the underlying task semantics. Experiments show that frontier models remain unreliable under these interactions, with substantial variation in task success and consistency across user conditions. Beyond evaluation, we use Tracer-generated data to fine-tune models across scales, resulting in consistent improvements on -Emotion-Bench and strong generalization to unseen domains, external benchmarks, and different user simulators. Overall, -Emotion-Bench and Tracer provide a scalable framework for evaluating agent reliability under diverse user behaviors and improving tool-calling agents through transferable supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.