acceptodds
Under review as a conference paper at ICLR 2027

When We Evaluate AI Scientists, What Are We Evaluating ?

Abstract

The highly complex and multifaceted nature of scientific research makes it particularly challenging to accurately assess AI agents in research workflows. Specifically, frontier AI trained in reward-seeking environment tend to learn shortcuts that achieve higher scores without improving the underlying science. We term this failure mode evaluation hacking where AI might learn to game proxy metric, leverage memorized data or look for easy answers online when presented with highly challenging tasks. In this study, we first present a systematic taxonomy on how evaluation hacking could occur and demonstrate their failure case on MLR-Bench, ResearchClawBench, and SciAgentGym. We find that susceptibility differs substantially across models and scientific domains. We observed that score pressure induce less faithful reports with forged numbers and sometimes in direct contradiction to objective evidence. These strategic misrepresentation can occur even when agents do not explicitly describe a plan to game the evaluation metric. Adding rules that account for evidence support can reduce, but not eliminate such harmful behavior across all 3 scenarios. In conclusion, this work seeks to serve as a pilot study that uncovers the covert failure mode of eval hacking in many AI for Research and RSI evaluation, even though no graders are broken and no questions are factually incorrect, frontier agents have shown increasingly stronger capability to infer and game eval metric without breaking/altering any code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.