ENURI: Evidence Negotiation Using Reasoning over Images
Abstract
A language-model agent acting on visual evidence must notice the evidence that matters (Detection) and weigh it correctly (Valuation), but outcome-level evaluations score only the final result and cannot tell which of the two failed. We introduce ENURI, a bilateral price-negotiation benchmark that extends the diagnostic TERMS-Bench environment with an image channel in which an item's defects appear only in its photo and carry a known ground-truth price impact. Placing defects on a visibility × alignment grid (Obvious/Subtle × Aligned/Misaligned) makes the correct response to evidence verifiable, so every failure can be attributed to one stage or the other. We evaluate 13 frontier, open-weight, and sub-frontier LLM/VLM agents and three rule-based baselines over 6,400 episodes, with rankings stable across four price-impact mappings. The two stages fail along different axes, and neither failure is visible in the outcome score. Detection recall falls from 0.551 to 0.278 on Subtle defects for all 13 agents, whereas Valuation accuracy holds between 0.92 and 0.99 in three cells of the grid and falls to 0.73 in Subtle-Misaligned alone. Claude Opus 4.6 captures the most surplus of any agent (normalized surplus 0.593) yet shows the largest Detection drop, from 0.612 recall on Obvious defects to 0.217 on Subtle ones, while GPT-4o-mini cites most precisely yet scores the weakest overall. ENURI thus shows not only which agents fail but which stage fails. We release the benchmark, dataset, and per-episode traces.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.