acceptodds
Under review as a conference paper at ICLR 2027

RTD-Eval and LongChemBench : Dual-View Evaluation of Long-Horizon Agent Tasks

Abstract

Recent progress in foundation models has enabled agents to use tools, interact with external environments, and pursue user-specified goals. However, their evaluation remains difficult because short-form benchmarks often fail to capture the multi-stage structure and verification demands of real-world workflows. We introduce LongChemBench, a long-horizon benchmark for computational chemistry, together with Report-Trajectory Dual-View Evaluation (RTD-Eval), a general grading framework for evaluating long-horizon agent tasks. RTD-Eval evaluates both the completeness and accuracy of the final report and the trace-report consistency. We instantiate this framework with 50 long-horizon open-ended computational chemistry tasks. When graded only on the final report, Qwen3.8-Max achieves the highest score, suggesting that it can produce comparatively complete and convincing answers. In the trace-report consistency view, most models exhibit substantial inconsistency, indicating that their final reports often include claims, values, or validation steps that are not fully supported by the actual execution process. This discrepancy highlights the limitation of final-report-only evaluation and motivates the need to evaluate both what an agent reports and how it arrived there.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.