Do AI Agents Have Research Taste? SENSE-Bench for Anomaly-Driven Scientific Exploration
Abstract
Existing AI research systems and benchmarks largely start from predefined ideas or research goals, while scientific inquiry can also begin with unexpected observations. Pursuing such observations requires research taste to judge which anomalies warrant investigation, a capability that remains insufficiently evaluated in current agents. We introduce SENSE-Bench (**S**cientific **E**xploration through **N**oticing, **S**electing, and **E**xperimenting), a benchmark for evaluating whether AI agents can recognize anomalies, judge which warrant investigation, and pursue experiments that support mechanism discovery. Across scientific domains, agents continue ongoing scientific investigations, identify unexpected findings, and decide whether and how to investigate them. We construct executable environments with fictional mechanisms inspired by real scientific phenomena, alongside noise and anomaly-free controls, providing known ground truth for evaluating agents’ judgments and discoveries. The evaluation measures anomaly detection, research taste, experimental support for proposed explanations, and conclusion correctness separately. The benchmark comprises 180 cases across nine families, and available results show stronger anomaly detection and mechanism discovery in GPT-6 Astra alongside weak research taste. These findings highlight the need to improve both the selection of promising anomalies and the experiments that turn observations into supported mechanism explanations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.