acceptodds
Under review as a conference paper at ICLR 2027

GUI-HITO: Benchmarking History-Induced Tool Overuse in Multimodal Agents

Abstract

Knowing the answer does not guarantee that a computer-use agent will answer directly. We show that prior tool use can bias agents toward further action even when the current request requires no interaction—a failure we call history-induced tool overuse. We introduce GUI-HITO, a paired benchmark that probes this answer–act boundary with English and Chinese questions across 33 history conditions. Comparisons hold the current request, screenshot, tool interface, and system prompt fixed while matching session counts and screenshot sequences across tool-bearing and action-free histories. On a shared, answerability-calibrated test panel, all five evaluated multimodal endpoints exhibit higher first-response action-attempt rates after tool-bearing histories, with increases of 29.0–53.8 percentage points and significance after multiplicity correction. Actions are scored without execution. The effect survives replaying prior actions as plain text, showing that native tool-call formatting is not necessary; four endpoints also exhibit increased attempts when the current screenshot already contains the answer. An answer-first prompt gate reduces attempts under tool-bearing histories from 60.6% to 16.2% on a development panel while leaving direct-answer accuracy unchanged. These findings expose a gap between execution competence and action restraint: reliable agents must recognize when to answer, even after a conversation full of actions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.