acceptodds
Under review as a conference paper at ICLR 2027

Obs2Rule: Benchmarking the Rule Discovery Ability of LLM Agents

Abstract

For agents aiming to conduct scientific discovery autonomously, uncovering a new law requires designing experiments, seeing through complex experimental observations, and finally extracting the regularity that produced them. We formalize this process as rule discovery, comprising two stages: experiment design, deciding which observations to collect and how to collect them systematically, and rule induction, compressing the accumulated evidence into a general statement that also holds for cases never observed. To evaluate both stages jointly, we propose Obs2Rule, a text-based simulation benchmark of 32 manually designed environments spanning 12 scenarios. Each environment offers no goal to achieve and no reward; the agent must autonomously design and carry out experiments to uncover the latent rule that governs the environment, and express the induced rule as executable Python code, which is then verified against manually crafted test cases. This design decouples rule discovery from the task completion measured by most existing scientific discovery benchmarks, and provides an objective, generalization-sensitive criterion for grading induced rules. Evaluating four LLM-driven agents, we find that the best agent passes only 66% of environments, and just 53% of those requiring active intervention. Further analysis shows that current agents can articulate rules once given adequate evidence, but are weak at designing and executing the systematic experiments needed to obtain it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.