Predicting User-Reported Problems from Coding-Agent Text
Abstract
Can a coding agent's prose help predict whether the user's next message will report a problem? We study 34,046 work-blocks from real coding-agent sessions and predict next-user reports that something is broken, erroring, or not working. These reports are only weakly associated with detected execution errors, making them a distinct feedback target. Adding within-block prose to structured context, developer history, and work statistics raises average precision from 0.211 to 0.258. The additive gain persists after excluding blocks with specified lexical cues. In separate user and temporal holdouts, prose alone also outperforms the feature baseline. At a 10% review budget, the combined model captures 35.5% of subsequent reports versus 30.2% for the baseline. The results establish prose as an incremental signal for offline review allocation, without claiming code correctness or intervention benefit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.