acceptodds
Under review as a conference paper at ICLR 2027

HLL: Can Agents Cross Humanity’s Last Line of Verification?

Abstract

Multimodal agents are increasingly expected to operate graphical interfaces on behalf of users, yet real workflows often contain verification gates designed to constrain automation. Existing GUI-agent benchmarks emphasize navigation and task completion, while such verification steps are omitted. This leaves an important capability under-examined: whether an agent can couple visual understanding with grounded interaction, state tracking, and process-consistent execution under controlled verification constraints. To address this gap, we introduce **HLL**, a controlled benchmark for interactive CAPTCHA-style verification in an authorized local simulator. HLL spans ten task families and factorizes evaluation along three dimensions: intrinsic task difficulty, webpage distraction, and trace-conditioned interaction validation. Agents must produce semantically correct answers through grounded GUI actions, while dynamic variants additionally validate simulator-observable evidence. We evaluate eight multimodal agents in a closed-loop Android/Chrome environment using a common Mobile-Agent-v2 harness. Within-run audits show that correct submissions can still fail the configured dynamic rules. Supplementary controls also identify legal strategies rejected by these rules. A repair study recovers most tested geometric rejections through trace editing. HLL measures simulator-level process compliance, not behavioral authenticity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.