acceptodds
Under review as a conference paper at ICLR 2027

Evaluating Agents in Live, User-Facing Tasks

Abstract

AI agents are increasingly expected to complete complex, real-world tasks on the live web, where success depends not only on open-ended reasoning but also on information that changes over time. Products fluctuate in price and availability, listings appear and disappear, deadlines pass, and policies are revised. Consequently, many web tasks have dynamic, open outcome spaces: they admit multiple substantively different valid outcomes, and the set of valid outcomes itself evolves with the state of the web. Existing benchmarks typically simplify one or both dimensions, or leave both to a single judge. They may allow agents to choose their own solution paths while evaluating against a pre-authored answer or canonical final state, or preserve reproducibility by freezing web content and avoiding time-sensitive tasks, or grade live-web outcomes with one LLM judge that handles constraints, evidence, and quality together. We introduce LIVEOUTCOME, a benchmark of 78 human-authored, user-facing tasks for evaluating agents against the current state of the web. Its tasks span research, comparison, planning, decision support, data collection, and artifact creation, with outcomes that may change across days, weeks, or months. Each task is paired with a hidden, property-based success specification that remains stable over time without enumerating a fixed answer. LIVEOUTCOME evaluates submissions through deterministic checks of explicit constraints, live retrieval and verification of source-dependent claims against the pages the agent cites, and rubric-based assessment of the resulting deliverable that incorporates the verified findings. We benchmark current agents and validate the evaluation protocol using human judgments of both LLM stages, repeated evaluation under fixed evidence, and controlled grounding failures with unmodified controls. Across 13 agent–model configurations, the strongest system scores 80.6 and passes 61.5% of tasks, and most failures stem from incomplete or incor- rect content rather than missing evidence. By separating stable success criteria from a changing set of web-grounded outcomes, LIVEOUTCOME enables rigorous evaluation of whether agents can solve open-ended tasks as the web exists at execution time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.