acceptodds
Under review as a conference paper at ICLR 2027

Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case

Abstract

Accident, defect, and outage investigations end with a decision that ordinary question answering never faces: whether the evidence collected so far is enough to close the case. We study this decision for LLM investigators. Given a case brief and an index of evidence items, the model requests evidence, revises its hypotheses, and either closes the case with a conclusion grounded in what it read, or leaves it open and names what is missing. This judgment does not come with capability: an untrained 9B model overstates its evidence in 97% of its answers, and a frontier model that identifies the right cause in 84% of cases still overstates in 91% and closes 17 of the 41 cases whose official finding is “cause undetermined”. Measuring the judgment is itself non-trivial, because the source of a case largely predicts its label: a rule that looks only at the source reaches 83.0 balanced closure accuracy. We therefore evaluate closure with three tests: closure accuracy, reported against the source-only rule and within each source; evidence dependence, which removes the grounds of a conclusion and checks whether the model stops closing; and conclusion and gap quality, a judged checklist of what the model asserts and what it says is missing. We build Nautil, 731 audited investigation cases from aviation, rail, maritime, chemical-safety and vehicle-defect reports and from production server incidents, with teacher trajectories, an out-of-distribution test set and counterfactual evidence versions. Fine-tuning a 9B model on these trajectories makes its closures follow the evidence: removing the grounds lowers its closure rate by 26 points (base model: 6; frontier model: 7), overstatement falls from 97% to 35%, and conclusions that are both correct and not overstated rise from 3% to 43% (frontier model: 9%). Reinforcement learning that rewards only the correctness of the closure decision then raises balanced closure accuracy from 69.2 to 83.3, on par with the teacher, at a measurable cost in evidence dependence. We release the models, the dataset and the evaluation suite.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.