Arxidea: An Agentic System for Personalized Research Ideation and Evidence-Backed Verification
Abstract
Language-model systems can propose research ideas at scale, but they leave two gaps for individual researchers. The proposals are rarely conditioned on the re- searcher’s own publication record or on papers appearing that week, and the result- ing ideas are typically handed over without verification. Existing systems often evaluate personalization using simulated users or model-based judges rather than the intended researcher. We present ARXIDEA, a personalized, human-in-the- loop agentic system that closes these gaps. Its four stages are Recommendation, Ideation, Desk Check, and Experiment Check. Recommendation ranks the week’s relevant papers using a profile derived from the researcher’s publications. Ideation proposes ideas from them. The two checks form a two-level idea checker that re- turns an evidence-backed feasibility verdict. We evaluate relevance directly with the profiled researcher. We construct a week-sized evaluation pool by sampling 2,574 new papers uniformly across twelve weeks of arXiv, and the researcher grades every paper for relevance. The resulting pipeline returns 100 papers of which 84 carry the researcher’s top relevance grade, spread across subject areas. Adding a reranker over the retrieved candidates reduces recommendation quality according to the researcher’s grades. A frontier agent (Claude Opus 5.5) that reads the week-sized sample in place of retrieval only matches it while spending 15 to 29 million tokens per week. Retrieval and ranking therefore operate without a gen- erative model. In the first level, the Desk Check, every candidate idea undergoes premise, theory, and novelty checks, after which it is retained with supporting ev- idence or eliminated with a recorded reason. The idea the researcher selects then enters the second level, the Experiment Check, a bounded experiment–review loop whose persistent memory is a versioned research specification. The loop records its verdict and hands surviving ideas back to the researcher with findings and next steps. We report three runs—one that passes, one that fails, and one that needs the researcher’s input—that expose the checker’s verdicts and supporting evidence. We will release the code, prompts, and researcher-provided relevance grades.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.