TIGER: Learning from Researcher Taste to Generate Tailored Research Ideas
Abstract
Large language models (LLMs) can generate research ideas, but their outputs often fail to match an individual researcher’s taste. We introduce Tailored Idea GEnerator for Research (TIGER), a framework that learns individual preferences to guide idea generation. Using 400 real-world research requests and 2,000 candidate ideas, we collect ratings from three LLM researchers and two PhD researchers. For each researcher, we train a separate Qwen3.8-27B reward model (RM) on their annotations, then fix its parameters and use its scores to optimize a Qwen3-8B ideation policy through reinforcement learning. Neutral rewriting and auxiliary quality checks guide optimization alongside the preference reward. In pairwise comparisons where each judge chooses between an idea generated by its corresponding TIGER-8B policy and one generated by GPT-4.1, the three LLM researchers select the TIGER-8B idea in 56.6%–68.4% of 76 comparisons per researcher. PhD 1 and PhD 2 select the TIGER-8B idea in 19 of 30 pairs (63.3%) and 17 of 30 pairs (56.7%), respectively. All six directed cross-policy comparisons likewise favor the judge’s matched TIGER-8B policy over an off-target policy, supporting the use of individual preference feedback for tailored research ideation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.