acceptodds
Under review as a conference paper at ICLR 2027

Do Research Agents Improve for the Right Reasons?

Abstract

Autonomous research agents are typically evaluated by whether an intervention improves a final metric. This conflates mechanistic understanding with feedback search, lucky exploration, and trial-and-error. We introduce a controlled benchmark for separating these possibilities in neural quantum-error-correction research. The environment exposes progressively richer evidence—from scalar error rates and syndromes to latent physical fault traces and counterfactual replays—and requires the agent to submit a machine-checkable diagnostic record before selecting a training-data intervention. The diagnostic is scored independently on held-out physical mechanisms, while its downstream decision relevance is tested by controlled record edits that hold the agent, evidence, execution budget, and decoder stack fixed. Across preregistered studies, outcome-only feedback improves decoder performance, demonstrating that endpoint gains can arise without mechanistic evidence. Counterfactual failure evidence improves intervention quality even without feedback search, while corrupting the submitted diagnosis substantially reduces utility; a serialization placebo has negligible effect. However, transfer across known physical shifts is not robust, revealing that local decision relevance does not by itself establish reusable mechanistic understanding. Our results provide a causal decomposition of research-agent improvement into evidence use, diagnosis, search, and intervention design, and turn “improvement for the right reasons” from an outcome-only claim into a falsifiable empirical question.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.