CLESO: Criterion-guided Laminar Evidence Sequence Organization for Specific Domain Budgeted Probative Evidence Selection
Abstract
Evidence selection in domain-specific question answering aims to identify a compact set of text units from a large-scale corpus that jointly supports an answer. Building on high-recall candidate sequences produced by relevance ranking, existing systems improve evidence selection through instructions, iterative reasoning, or set-level objectives. However, evidence set selection in domain-specific QA must locate support that may be dispersed across long candidate sequences and assess whether a candidate set jointly enables a reliable answer without mixing incompatible entity conditions or scopes. Repeatedly assessing alternative sets with an LLM is costly. We propose CLESO, a structural sequence distillation framework that learns laminar candidate sequences to prioritize access to potential support under a limited observation budget before downstream joint selection. During supervision, ontology-grounded criteria provide signal for candidate qualifications and cross-evidence dependencies, and a Laminar compiler turns these signals into local ordering, head-admission, and boundary-separation targets. At inference, a lightweight Transformer refiner predicts residual scores from inference-available features, and a fixed decoder recovers intervals from the resulting sequence. The recovered intervals guide budgeted local observation before LLM-based package selection, allowing the learned structure to guide evidence access without executing criterion search at inference On BioASQ, QASPER, and LegalBench-RAG, CLESO-CP gains an average of 1.45 evidence-F1 points over the strongest baseline, with a 3.25k-token observation cap on queues averaging 1M tokens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.