acceptodds
Under review as a conference paper at ICLR 2027

Told Who Could Know: False Access Labels Mislead LLM Aggregators

Abstract

A large language model (LLM) can act as an aggregator: it reads other models' answers and returns a final answer. We test whether access labels, which claim who received information needed to answer, influence this choice. Four LLMs answer from assigned options or documents, so we know which labels are true. We swap only the labels, keeping questions and responses fixed. Aggregators see assigned inputs only through responses. Across five aggregators on multiple-choice questions (MMLU-Pro and SuperGPQA) and document questions (2WikiMultiHopQA and MuSiQue), replacing true labels with false ones shifts answers toward the falsely labelled model and lowers accuracy. On MMLU-Pro science, technology, engineering, and mathematics (STEM) questions, accuracy with full responses is 20.8–37.2 percentage points lower under false labels than without labels. Adding accompanying text to final answers narrows the true-versus-false accuracy gap on STEM and both document benchmarks. On STEM, we tell aggregators that the positively labelled model almost never received the correct option. All five aggregators select that model's answer less often after disclosure. Three still select that model's answer 15.2–29.2 points more often than without labels, and all five remain less accurate. False access labels can redirect aggregation even when their unreliability is stated.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.