acceptodds
Under review as a conference paper at ICLR 2027

IdeaDec: Benchmarking Long-Horizon Scientific Idea Decoding in AI Research Agents

Abstract

AI research agents increasingly participate in literature analysis, hypothesis formulation, experiment design, and proposal development, making the coherence of their evolving ideas consequential to downstream research. Yet existing evaluations largely score final outputs or isolated stages and therefore do not test whether agents sustain valid scientific commitments as evidence changes. We introduce IdeaDec, a benchmark for long-horizon scientific idea decoding under controlled evidence release. IdeaDec operationalizes this problem through 137 trajectories of 7–9 ordered checkpoints, each tracking a five-component state: hypothesis, method, metric, judgment, and plan. Three disjoint subsets evaluate dependency maintenance, conflict-aware revision, and dynamic recovery, while complementary metrics assess structural completion, success on a pre-specified target operation, and five-gate whole-trajectory semantic success. Across 18 fixed model configurations in the State-Aware setting, mean structural completion is 89.05% but mean semantic success is 51.95%, a gap of 37.10 percentage points; similar aggregate scores can also conceal distinct capability profiles. Recovery has the lowest observed subset rate for every configuration under the fixed benchmark distribution. These findings motivate evaluating scientific agents through state evolution rather than format compliance or endpoint outputs alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.