UnitAlign: Citation Repair that Preserves Supported Claims
Abstract
A citation-grounded answer can contain one unsupported sentence amid several that already have valid evidence. Regenerating the answer reopens those verified claims, whereas citation scoring alone leaves the failed claim untouched. UnitAlign turns the sentence and its cited spans into a shared interface for human inspection, verifier decisions, and corrective action. A budgeted packet compiler removes redundant evidence; sentence-level provenance then permits a calibrated external verifier to trigger focused retrieval and rewrite only the failed unit. On BioMed-180, nine Llama-3-70B-Instruct systems receive the same corpus, citation schema, and 8,192-token visible-evidence allowance. In 360-pair expert audits per system, UnitAlign reaches 94.7% supported-sentence rate and 93.1% citation precision—gains of 7.6 and 11.7 points over matched-stack RARR—and exceeds the strongest faithful native citation system by 2.4 and 1.7 points. A fixed-state intervention sharing initial answers, citation bindings, verifier decisions, and triggers shows why: the local policy improves support on triggered units by +6.1 [+1.4, +10.9] points, preserves +44.2 [+38.1, +50.5] more previously supported content, and saves 7.01 [6.42, 7.63] s relative to whole-answer regeneration. Failure locality—unsupported-unit density, concentration, and discourse coupling—also reveals the escalation boundary: local repair leads by 8.9 support points for sparse, weakly coupled failures, while answer-wide repair leads by 5.8 for dense, strongly coupled failures. Complete public audits and Qwen and Mixtral generators preserve the comparative advantage; verifier swaps retain 93.2–95.2% human SSR. Fine-grained attribution can therefore govern corrective action, not merely document it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.