Lazy Literature Review: Quantifying Sloppiness in LLM Research Agents
Abstract
LLM agents are increasingly trusted to do research: this includes compiling evidence from a corpus of literature and making recommendations or decisions about further steps. A recognised weakness of such agents is sloppiness, work that is left incomplete but reported as finished. In evidence synthesis this is hard to notice, because a fluent, well-cited answer looks the same whether or not the evidence supports it. We set out to quantify this sloppiness, and chose one failure mode that can be verified exactly: treating results as comparable when they are not. We build synthetic corpora of literature in which the studies bearing on a contested claim measure different things, so their numbers cannot be combined and the right conclusion is known by construction. We score whether an agent reaches a conclusion without noticing the incompatibility and, separately, whether it uses the incompatible numbers in its conclusion at all. Across twelve frontier models, most notice the problem and pool the numbers anyway: some name the incompatibility in over 90% of answers and still conflate in a third to a half. Only the strongest model, GPT-6 Astra, avoids the error on our synthetic literatures. On real arXiv papers whose results rest on incompatible protocols, every model, Astra included, still ranks by the incompatible numbers. Beyond the results, we propose the construction and scoring methodology as a general recipe for studies of this kind, and release the protocol and example corpora.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.