What Does LLM-Judged Novelty Reward? Controlled Experiments on Context, Workflow, and Measurement Validity
Abstract
Large language model (LLM) research-ideation systems are increasingly evaluated with LLM judges of novelty, yet what those judgments reward is unclear. We vary paper contexts, generation instructions, and revision guidance while holding the research question and main rubric fixed. Judges see only the question and idea texts, in both display orders. The rubric rewards substantive departures from known work. An aligned generation instruction requiring a component beyond combining known methods increased preference across two generator families and two judge families. Giving this instruction only to GPT-5 mini reversed the preference for GPT-5. When the generator was asked to plan verification before revising a draft, the judge more often preferred the version revised without that instruction, even though the proposed verification was never carried out. Retrieving prior work and giving the generator feedback on possible overlap before each review round produced ideas the judge preferred to those from general review. To assess how these judgments relate to scientific novelty, we also examined consistency and agreement with existing expert ratings. Reversing display order changed the judge's preferred idea in nearly one fifth of pairs answered in both orders, and its choices only partly matched expert preferences. The judge's preference alone does not establish that the favored ideas make more original scientific contributions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.