Measuring and Reducing Over-Reliance on Convenient Evidence in LLM Agents
Abstract
LLM agents often over-rely on easily inferable information, committing to convenient sources or hints without verifying them. This gullibility is often unrepresented in benchmarks where logic steps are well defined, but abundant in everyday life. We attempt to measure this phenomenon by introducing GulliBench: a set of 50 simple data tasks, each with a planted wrong shortcut and a correct answer that must be derived from the underlying records. The tasks are designed to be well within the coding and reasoning capabilities of frontier models, save for the over-reliance traps. Of the 23 models we tested, even the best one fails on 39% of its attempts, and the other 22 fail on more than half of theirs. We conjecture that this gullibility reflects a behaviour that general capability does not reliably produce, and that it can be learned. By post-training a 35B-parameter LLM with RL on a different set of tasks that expose these contradictions, we improve its accuracy on the benchmark from 3% to 31%. On third-party data-analysis and terminal benchmarks the post-trained model checks its work more often, but this is barely reflected in their scores, and we show through case studies why such benchmarks rarely reward it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.