Same Paper, Different Decision: A Controlled Diagnostic of Task-Conditioned Scientific Evidence Readiness
Abstract
Scientific papers are not intrinsically ready or unready for information extraction: their usability depends on the evidence required by the downstream task. We formulate task-conditioned scientific evidence readiness, in which a system maps a paper–task pair to one of three operational actions: READY, REVIEW_REQUIRED, or ABSTAIN. To isolate task conditioning from document variation, we construct a paired diagnostic set from eight Ti–6Al–4V laser powder bed fusion papers and three extraction tasks that differ only in their material-subtype requirement, yielding 24 correlated evaluations. Holding the paper fixed, the frozen AI-assisted author-reference action changes across tasks for six of eight papers. We compare four frozen prompting and control configurations. Direct prompting matches the reference action on 16/24 pairs, while the full structured configuration matches 14/24, indicating that additional structure does not consistently improve decisions in this testbed. An evidence-grounded configuration makes no false READY decisions among its 18 valid outputs but has execution failures on 6/24 pairs, exposing a safety–coverage tradeoff. Error analysis further shows that an output can record a task-relevant evidence conflict yet select an inconsistent operational action. These errors motivate a stage-wise diagnostic perspective separating evidence-state recognition from task-dependent action mapping; the stages are not independently evaluated here. Our results motivate pair-conditioned evaluation for scientific extraction systems whose outputs trigger downstream use or human review.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.