acceptodds
Under review as a conference paper at ICLR 2027

From Black-Box Judgments to Auditable Graphs: Structured Idea Generation and Evaluation via Scientific Knowledge Graphs

Abstract

Recent advances in large language models (LLMs) have fueled growing interest in autonomous research-idea generation. Now that many systems can propose ideas, candidates are no longer hard to come by; what is hard is deciding which of them justify the cost of implementation. This shifts attention to idea evaluators and raises a more fundamental question: how should these evaluators themselves be evaluated? Existing benchmarks rely on two unreliable signals: black-box LLM judgments that are unauditable and prompt-sensitive, and proxy labels, such as peer-review scores, publication time, and expert annotations, that are validated against human agreement rather than against actual implementation outcomes. We present SciGraph, a graph-based pipeline for generating and evaluating research ideas. SciGraph decomposes historical literature into meta, mechanism, and protocol nodes and connects them through citation, mechanism-combination, and protocol-result comparison relations. Candidate ideas are randomly sampled from a predefined space and rendered as text. Two fixed policies—mainstream and novelty-oriented—then select candidates for execution, differing only in how they weight three graph-derived priors. To evaluate an idea, SciGraph maps its text back to graph nodes and scores it using the same priors. The pipeline also yields SciGraph-Probe, a probe set for assessing idea evaluators. It contains auditable graph representations of 184 papers across 10 agent-related domains, comprising 184 meta nodes, 490 mechanism nodes, and 727 protocol records connected by more than 1,700 edges. It also includes 100 randomly sampled candidate ideas that undergo code synthesis and execution, producing effectiveness labels withheld from all evaluators, together with external bibliometric reference orders for novelty and importance. We conduct two separate experiments. On three idea-generation benchmarks, using the context provided by each benchmark, the two policies show the expected differences. On SciGraph-Probe, SciGraph ranks first on all three dimensions—effectiveness, novelty, and importance—and every score it produces is traceable to individual nodes and edges.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.