acceptodds
Under review as a conference paper at ICLR 2027

TAFER: Auditing Selective Repair Under Evidence Shift

Abstract

Language-model outputs can match task targets without being supported by the available evidence. We introduce TAFER, a framework for auditing selective repair under evidence shift. TAFER separates answer-text correctness, answer-or-abstain decisions, and support assessments tied to explicit reference rules. In public fixed-pool comparisons, its deterministic selector chooses the lowest recorded-cost candidate that passes structural citation checks, or abstains when none passes. Shared-pool baselines distinguish feasibility filtering from ranking, while matched evidence interventions examine how removing cited or uncited evidence changes answering behavior. Finite-pool accounting separates candidate unavailability, certificate exclusion, and selection error; risk-coverage analysis makes abstention costs explicit. Results are heterogeneous: changes in a support proxy do not alone establish better task utility, and filtering gains need not imply ranking gains. Behavioral diagnostics distinguish increased abstention from continued answering after evidence removal. Natural-language intervention support uses operational verifier judgments, while candidate-pool support uses a gold citation-ID relation. Neither is treated as human support truth. These analyses characterize selective repair within tested protocols without conflating correctness, support, and abstention.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.