Beyond Context Selection: Consensus over Responses to Alternative Contexts
Abstract
Large language models can produce different predictions depending on how inference-time context is composed or applied. Existing approaches have largely addressed this variation by deciding which context to provide before prediction. We instead ask whether the target model's responses to alternative contexts can be used to adapt the final decision to each query. We propose CHORUS, a test-time consensus framework that complements context selection by jointly using responses from alternative ways of composing or applying the supplied context, without training a separate selector. CHORUS combines answers that are repeatedly supported across alternative contexts with answers that receive strong support from a single context. The same principle applies to different demonstration compositions in in-context learning and to different reasoning plans applied to fixed skill documents. Across seven LLMs, 12 classification tasks, six structured and short-form generation tasks, and four skill-augmented reasoning domains, CHORUS shows strong performance against the evaluated selection and consistency baselines. Further analyses show that their relative strengths change with the amount of available context and across tasks, and that adapting consensus decisions to the current query improves prediction. These results support using responses to alternative contexts as test-time evidence beyond context selection alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.