Predictable Labels, Fragile Rankings: An Empirical Study of LLM Evaluator Procedures
Abstract
Automated hypothesis generation creates a selection problem: which candidates should proceed to expert review, simulation, or experiment? Benchmarks assess evaluators through checkable labels, but a label’s predictability and an evaluator’s ability to recover it are different properties. We examine both on ResearchBench, whose fixed candidate pools link one hypothesis to a designated source study. Across 1,313 eligible items, a supervised cross-fitted candidate-only classifier achieves 98.40% sixteen-way Hit@1 without the background or research question. On the same pools, six zero-shot LLM evaluators average 55.79% source–alternative pairwise recovery when scoring candidates separately (Isolated), but 35.99% when generating their scores together (Vector), below the 50% pairwise reference. These endpoints diagnose label signal and procedure behavior under different training conditions; they are not a common ability ranking. Paired behavioral analysis shows that similar strict-win rates can conceal substantial exchanges among wins, ties, and losses. Controlled comparisons further show that replacing surrounding material can reduce recovery while preserving the target text and scalar output action, with effects that depend on the model and setting. With visible materials matched, it remains unresolved whether joint rather than separate scoring increases or reduces this loss. The evidence supports evaluating label meaning, candidate discrimination, and recovery in the intended context separately before treating benchmark performance as evidence of reliable scientific screening.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.