acceptodds
Under review as a conference paper at ICLR 2027

The Novelty Trap: When Repetition Helps LLM Research

Abstract

Novelty is a common objective in automated research, but forbidding repetition can reduce the value of what agents discover. We demonstrate this novelty trap on the Koijen–Levy earnings-announcement benchmark. An LLM proposes economic hypotheses, a second model measures them from event summaries, and hypothesis sets fitted in one quarter are evaluated in two later quarters. At matched proposal budgets, independent restarts improve incremental R² by 23% and 19% relative to a sequential no-repeat prompt (0.63 and 0.57 percentage points). The advantage holds on all six seeds in each quarter, including evaluation of the same runs on historical labels withheld until the analysis protocol was fixed. A memory-preserving prompt that permits repetition also improves value on every seed. We characterise the exploration–measurement tradeoff with a model separating errors in how hypotheses represent constructs from errors in their annotation. Under its assumptions, an informative prior and noisy measurements create a region in which repeated formulations of strong ideas outperform broader coverage. The model also predicts exactly when additional annotation can reverse that ranking. Planted-truth experiments illustrate the value of exploration when prior information is weak. Explanatory value and discovery counts also diverge: counting significant coefficients in joint regressions reverses the policy ranking. The design implication is to broaden construct coverage while preserving opportunities to refine useful mechanisms, and to evaluate both by the joint information they add.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.