More Than Can Be Said: A Benchmark and Framework for Pre-Question Scientific Ideation
Abstract
AI research agents have shown strong potential in automating literature search and manuscript refinement, yet most operate only after a research question has already been made explicit, assuming a clear and actionable initial input. Human research, however, often starts earlier, with an inarticulate sense of friction between intuition and existing explanations, before any question can be posed. We introduce InciteResearch, a multi-agent framework that helps a researcher externalize this pre-question friction into an inspectable, revisable profile, and then uses that profile to reframe the underlying problem before drafting a proposal. Concretely, InciteResearch (1) elicits a five-dimensional researcher-profile representation, comprising an explicitly logged friction point together with motivation, constraints, preference, and a refined topic, through structured multi-turn dialogue; (2) identifies a hidden assumption underlying the initial framing by jointly scoring candidate assumptions on removability and reframing potential, and produces a reframed problem statement with a seven-stage causal derivation trace linking assumption, insight, claim, prediction, and method; and (3) checks whether each proposed method component is a logically required consequence of that derivation rather than a plausible add-on. We introduce TF-Bench to evaluate InciteResearch. TF-Bench is a benchmark for tacit-to-explicit research assistance spanning prediction, discovery, attribution, causality four scientific modes and domain-related, domain-unrelated two ambiguity regimes, now comprising 120 inspiration examples. Relative to a matched AgentLaboratory and prompt-based baselines, InciteResearch improves mean novelty and impact from 3.617 / 3.810 and 3.695 / 3.827 to 4.221 / 4.368, respectively, showing consistent gains across all four domains, which are further corroborated by expert evaluations validating the reliability of the judges. These results support assumption-breaking dialogue as a measurable, specific contributor to proposal quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.