acceptodds
Under review as a conference paper at ICLR 2027

Are Research Agents Fooled by Clever Hans? A Benchmark for Proactive Pitfall Detection in Medical Autonomous Research

Abstract

Research agents built on language models now design experiments, run code, and draft claims across long research workflows. Along the way they can introduce methodological errors that make a result look strong for the wrong reason, and in medical AI such a result can travel into a deployment claim with clinical consequences. It remains unclear whether these agents catch such errors on their own or, like the audience of Clever Hans, accept an answer because it looks right. We introduce MAP-Eval, a benchmark for proactive pitfall detection in medical autonomous research. Each of its 189 medical-AI workspaces contains one methodology pitfall drawn from a 33-type taxonomy, and the agent is never told that a flaw exists. MAP-Eval holds each pitfall fixed across nine harnesses, from single-turn review and a tool-using ReAct loop to review states captured from three autonomous-research frameworks. Across ten frontier models, static review reaches 66–70% mean identification rate, the six framework states reach 41–50%, and a shared review checkpoint after data preparation holds every model between 21% and 34%. Tool access raises some models by up to 26 points and lowers others by up to 37. Further analyses suggest that the bottleneck is applying methodological knowledge inside a workflow, not acquiring it. The same models that name a pitfall when handed the whole workspace overlook it when the evidence is spread across files, when they must find it with tools, or when the current stage asks them to write the result up as a claim. MAP-Eval provides a controlled testbed for measuring whether agents apply methodological knowledge where the workflow needs it, and a foundation for autonomous research that catches its own errors before they reach a clinical claim.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.