RTD-Eval and LongChemBench : Dual-View Evaluation of Long-Horizon Agent Tasks
Abstract
Recent progress in foundation models has enabled agents to use tools, interact with external environments, and pursue user-specified goals. However, their evaluation remains difficult because short-form benchmarks often fail to capture the multi-stage structure and verification demands of real-world workflows. We introduce LongChemBench, a long-horizon benchmark for computational chemistry, together with Report-Trajectory Dual-View Evaluation (RTD-Eval), a general grading framework for evaluating long-horizon agent tasks. RTD-Eval evaluates both the completeness and accuracy of the final report and the trace-report consistency. We instantiate this framework with 50 long-horizon open-ended computational chemistry tasks. When graded only on the final report, Qwen3.8-Max achieves the highest score, suggesting that it can produce comparatively complete and convincing answers. In the trace-report consistency view, most models exhibit substantial inconsistency, indicating that their final reports often include claims, values, or validation steps that are not fully supported by the actual execution process. This discrepancy highlights the limitation of final-report-only evaluation and motivates the need to evaluate both what an agent reports and how it arrived there.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.