NicheLedger: Replacing References and Graders on Fixed Answers of Spatial Transcriptomics Agents
Abstract
Some scientific-agent benchmarks grade open-ended answers with a language model against a fixed reference. Replacing either can move a score on unchanged answers. Using NicheLedger, a spatial-transcriptomics agent benchmark, we replace each on fixed answers and track levels, orderings and test outcomes. On 36 niche-naming items under one grader, a frontier model's answers score 27.8% against clustering-derived references and 77.8% against clinician-consensus references, two physicians' hand grading gives 30.6% and 86.1%, and item-derangement controls score at most 8.3%. Of 18 gained items, seven clustering-derived references name a cell type no clinician named, three name no niche, and eight are wording, granularity or unresolved cases. One of 21 system-pair orderings reverses. The same model's 40 niche-naming answers score 30.0%, 47.5% and 20.0% under three grader states, two reached by names reporting one model, counting unparsable replies incorrect. Three claims each change outcome under some state. On BixBench, the ordering survives three graders and model-proposed references that two physicians prefer, which move accuracy by up to 4.5 and 1.4 points respectively. References and graders thus move levels and test outcomes while most orderings hold, and clinician consensus is not biological ground truth.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.