acceptodds
Under review as a conference paper at ICLR 2027

HTMLPresentBench: How Well Do Agents Present Task Results Through HTML?

Abstract

AI agents can carry out complex tasks, but their usefulness also depends on how effectively the results are communicated to users. Long outputs and execution traces often make it difficult for users to extract key findings and assess the evidence supporting them. We study how to evaluate reports generated from completed agent trajectories. These reports should faithfully convey important outcomes and limitations, while making results and supporting evidence easy to understand and find. We introduce an assessment framework for trajectory-grounded reporting that evaluates reports along three axes: content, visual quality, and information access. By grounding evaluation in the underlying trajectory, task evidence, and readers’ information needs rather than a single reference presentation, it accommodates diverse report designs while providing explicit, inspectable criteria. To instantiate and validate this assessment, we introduce HTMLPresentBench, containing 60 completed task trajectories across 13 domains, and evaluate eleven report generators that turn the same fixed trajectories into user-facing HTML reports. Our evaluation reveals tradeoffs across content, visual quality, and information access. HTML improves information access on average over Markdown, but can reduce content or visual quality, while design-oriented skills further improve access in some settings at the cost of other dimensions. Human evaluation further provides evidence that the full assessment pipeline better reflects readers’ judgments than a simple LLM-as-a-Judge baseline. Because trajectory-specific references are constructed once and reused across candidate reports, the framework supports scalable, controlled comparison of models and reporting methods on shared executions, and can be applied to new trajectories and tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.