Beyond Answer Preservation: Scientific Evidence Minimization for Private LLM Reasoning
Abstract
Large language models can support scientific reasoning, but using external models often requires disclosing private scientific context. Reducing this disclosure while retaining useful answers is not sufficient for grounded reasoning: response preservation and downstream utility do not by themselves guarantee that the released context contains sufficient evidence to support the answer. We formulate **Scientific Evidence Minimization**, which seeks a policy-compliant, low-disclosure context that remains sufficient for grounded scientific reasoning. We introduce **CEDAR**, a verification-guided framework that iteratively evaluates an external answer, diagnoses missing evidence, and selectively expands the released context within the disclosure policy. We evaluate CEDAR across multiple scientific domains, reasoning tasks, and disclosure policies, explicitly separating answer correctness from evidence sufficiency to characterize how much scientific information is actually necessary for grounded reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.