How Does Knowledge Organization Shape LLM-Based Scientific Ideation
Abstract
AI scientists powered by large language models (LLMs) draw on domain knowledge to propose and develop research ideas. Existing systems organize this knowledge differently, as retrieved papers, literature chains, or concept graphs. Comparing complete systems cannot reveal which organizational choice shapes the resulting ideas, because such comparisons change the delivered evidence and its arrangement at once. We introduce a controlled evaluation framework that separates these two factors: knowledge access (what evidence reaches the model under a fixed budget) and knowledge presentation (how the same evidence is arranged). Across two studies with three generators and two judges, we run more than 68,000 idea generations and 87,000 model evaluations. On ResearchBench, access methods change the evidence the model receives and, where that evidence differs most, the benchmark-matching score. Enlarging the candidate pool dilutes the relevant evidence admitted and does not improve matching. With the evidence held fixed, linear and relation-grouped presentations are practically equivalent in all four generator–judge pairings. In multi-step development built on CHIMERA, a recombination instruction raises the rate of explicit citation of knowledge from other research branches by 52–62 percentage points over refinement, whereas two access methods delivering similar evidence show no consistent difference. These results indicate that knowledge organization shapes the resulting ideas mainly through which evidence reaches the model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.