ProofCover: Auditing Proof Replacement under Answer-Preserving Updates
Abstract
Evaluating rule-based reasoning by final-answer accuracy is standard practice, even when models also return the facts and rules supporting their predictions. However, after the rule world changes, the correct answer may remain the same while the submitted support becomes invalid, causing answer-centric evaluation to award full credit to an invalid certificate. Our key insight is that this is a structural observability gap: answer-only scoring cannot distinguish a valid current-world certificate from an invalid one when both predict the same correct label. Based on this observation, we introduce ProofCover, an evaluation framework and diagnostic benchmark that makes support validity under answer-preserving updates directly measurable. Instead of scoring only the final answer, ProofCover enumerates the complete set of minimal proofs for each query, uses this proof set to construct updates that preserve the correct answer while selectively invalidating the model's submitted certificate, and then applies a deterministic current-world verifier to check whether the model's new certificate remains valid in the updated world. On Formal-500, proof-set-aware construction achieves 100% success on Repair, compared with 16.6%–43.4% for three simple selection baselines. Under Fresh Adaptive Repair, false passes occur in all eight evaluated models: among outputs accepted by answer-only scoring, 15.3%–48.1% (median 27.1%) carry certificates that fail current-world verification. This pattern replicates on RuleTaker-MP, a RuleTaker-derived multi-proof benchmark constructed from official test data, where false-pass rates span 6.2%–22.0% across the same eight models. These results show that answer-only scoring can substantially overestimate support reliability in the evaluated update settings, while highlighting a certificate-validity gap that final-answer accuracy does not measure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.