acceptodds
Under review as a conference paper at ICLR 2027

The Intent Gap: A Taxonomy of Real-User Failure Modes in Frontier AI Agents

Abstract

Public benchmarks for frontier AI agents (such as GAIA, OSWorld, and SWE-bench) are uniformly researcher-designed: tasks are framed in clean declarative specifications with fixed ground-truth answers. Real users in natural deployment do not interact with models this way. They omit implicit context, revise goals mid-task, and abandon conversations silently. Consequently, static benchmarks systematically fail to detect the intent gap: instances where a model produces an output that is literally compliant with the prompt but misses what the user actually intended. We introduce a scalable frustration-signal filtering pipeline (lexical filtering, semantic similarity gating, and model-based verification) to extract intent-gap failures from public interaction corpora. Applying this pipeline to a 50,000-conversation slice of WildChat-1M surfaces critical empirical dynamics: (1) 52.8% of real-user English interactions terminate after a single turn, highlighting widespread silent abandonment invisible to feedback-driven pipelines; (2) real-world repair markers differ drastically from academic literature, where blunt conversational cues outperform researcher-imagined phrases by  31x; and (3) standard safety-filtered judge models refuse to evaluate 62% of real-user candidate failures, exposing a structural blind spot in post-deployment evaluation pipelines. We formalize an 11-category taxonomy of real-user failure modes, triangulate observed failures against five documented production incidents, and release an open-source evaluation suite and verified corpus to bridge the evaluation-to-deployment gap.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.