Beyond Correctness: The Maximal Supported Answer
Abstract
Answer correctness is usually evaluated against a single target, even when the available evidence licenses several answers at different resolutions. A model that says only "in the twentieth century" despite decisive evidence for 1994 is correct but uninformative; a model that says 1994 from decade-level evidence is informative but unsupported. We formalize this tension through the maximal supported answer: the most specific answer whose denotation contains every world compatible with the supplied evidence. In support-complete answer lattices, this target is unique and moves monotonically toward specificity as evidence is refined. We instantiate the framework in MaxSupport, a paired evidence-ladder benchmark spanning time, geography, numerical intervals, entity taxonomies, and source attribution. We study LatticeCal as a calibrated non-oracle instantiation that estimates candidate-level support risk and chooses the deepest feasible node without accessing the denotational oracle at test time. In an author-constructed synthetic evaluation across five semantic spaces, LatticeCal reaches 64.5 exact-MSA with a 4.2% support-violation rate. CSS remains slightly stronger on geography and attribution, while HSC leads on taxonomy, showing that answer-space geometry determines which form of calibrated backoff works best. Paired additions and removals further reveal whether resolution tracks evidence rather than parametric recall, and claimwise risk allocation extends the framework to long responses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.