acceptodds
Under review as a conference paper at ICLR 2027

How Stable Are Answers to Unresolved Scientific Questions? Separating Repeat Variability from Input Sensitivity

Abstract

When a language model gives different explanations for the same unresolved scientific question, does the difference reflect a change in the input or variation across repeated generations? Without an accepted complete explanation, disagreement alone cannot distinguish these possibilities. We introduce Science80, a controlled study of 80 questions across five scientific domains, and compare evidence reordering, deletion, and addition against responses to identical prompts. GPT-6 Astra xhigh provides the main experiment, with same-question comparisons from Gemini and Claude. Under identical inputs, the question-weighted rate of substantive revision is already 58.3%, including changes in secondary mechanisms, scope, or scientific preference. Evidence reordering does not show a significant average increment beyond repetition, whereas adding a statement increases the substantive-revision rate by approximately 20 percentage points. This increment remains after adjustment for response length and other text features. Further analyses show that a change in the preferred explanation can accompany little change in reported confidence, while limited directional evidence constrains conclusions about position and conflict. Our results support evaluating input responsiveness relative to repeated-generation variability, while assessing explanatory stability, evidence responsiveness, and scientific correctness separately.rrectness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.