acceptodds
Under review as a conference paper at ICLR 2027

LiteratureRank: Bias-Reduced Leaderboards from Scientific Literature

Abstract

We introduce LiteratureRank, a literature-based approach for ranking methods in scientific domains that lack standardized leaderboards. The needed comparisons are often already published, but scores from different papers are not comparable, and most tables come from papers that propose one of the models they compare. We compare scores only within an evaluation cell, the scores that share a paper, a dataset and a metric, and expand each cell into pairwise wins and losses. To reduce author bias, self-report exclusion removes each paper's own proposed model from its tables during LLM-based extraction, so a model is ranked only from comparisons reported by other papers. A pair-weighted Bradley–Terry model turns these battles into scores with uncertainty intervals. We release 14 leaderboards and their comparison datasets, spanning general LLM and VLM capabilities and six scientific domains, which rank 1,082 models from 401k comparisons. A median of 74% of model pairs per board have non-overlapping 95% intervals, interval width shrinks as in the number of comparisons, and widening one board's collection window by a year leaves the order of its existing models nearly unchanged (Spearman ). The rankings can be updated as new papers appear with little additional computation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.