The Scientific Research Taste Benchmark (SciTaste-Bench): A Multi-Disciplinary Benchmark for Assessing Large Language Models' Judgment of Research Taste in Scientific Papers
Abstract
Whether large language models can assess the innovativeness of academic publications is important in literature reviews, research funding screening, and academic recommendations. We introduce the Scientific Research Taste Benchmark (SciTaste), a gold-standard benchmark covering six disciplines: physics, mathematics, chemistry, biology, medicine, and finance. SciTaste measures research taste, that is, the ability to judge the scientific value and potential of a paper, of which innovativeness is only one dimension among the seven assessed. SciTaste integrates metadata from NASA ADS, OpenAlex, and arXiv, aligns papers to field-specific classification systems, and evaluates each paper using the seven-dimensional Academic Taste Index. The weights for each dimension are derived from a bi-objective machine-learning derivation (five regressors, including random forests) and are shared across all six disciplines. This benchmark includes 300 selected papers (50 per discipline) and a backup pool of 5,700 papers (950 papers per field) covering 1996–2023. To allow external knowledge to influence model decisions, SciTaste supports four evaluation modes: pure_qa, native_search, our self-built structured tool-calling agent DimAgent (9 bibliometric tools in ADS/OpenAlex), and dimagent_v3_native. Model outputs are evaluated with two complementary schemes: one combines correlation with dimension-level accuracy, and the other is based on the WAPE relative error. In the evaluation of ten state-of-the-art large language models, the tool-based structured evaluation generally outperformed the toolless and raw web search modes, but the improvement was model-dependent. Physics is the easiest subject, while chemistry, biology, and medicine are still more challenging. By measuring how well large language models judge the value potential of academic work across fields, SciTaste supplies the value-judgment substrate that auto-research systems need to screen, rank, and prioritize candidate ideas. We release the dataset, the evaluation harness, and the scoring code to support reproducible research on scientific-innovation assessment; all data and code will be made publicly available on Hugging Face.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.