acceptodds
Under review as a conference paper at ICLR 2027

Recoverable Proofs: Reverse-Semantic Preference Training for Logical Reasoning

Abstract

Proof-first prompting encourages large language models to solve logical-reasoning tasks by generating intermediate proofs, rationales, or symbolic-style traces before producing a final answer. However, such traces are not necessarily faithful carriers of the original problem semantics: a model may predict the correct label while omitting key conditions, confusing entities or operators, or relying on shortcuts that are difficult to recover from the trace itself. We call this the proof-trace recoverability gap. We propose Reverse-Semantic Preference Training (RSPT), a train-time plug-in method for proof-first logical reasoning. For each problem, RSPT samples candidate proof traces, extracts deterministic semantic slots from the original problem, constructs corrupted slot alternatives, and uses the frozen backbone to score whether the generated trace makes the gold slots more recoverable than the corrupted ones. This recoverability signal is then used to construct same-question preference pairs for standard SimPO/LoRA adapter training. At inference time, the trained adapter produces one deterministic output under an existing proof-first scaffold, without reward scoring, reranking, majority voting, or an additional verifier. Across FOLIO and ProofWriter, RSPT improves all 9 of 9 complete parser-safe same-scaffold settings under the same backbone, parser, decoding budget, and greedy one-output protocol, with gains up to +5.23 points on FOLIO and +10.78 points on ProofWriter. Under a fixed global Symbolic-aided-CoT deployment scaffold, RSPT-Scaffold@1 reaches 69.93% ± 0.57 on FOLIO and 69.28% ± 2.74 on ProofWriter, outperforming the strongest frozen parser-safe baselines in our unified setting. We separate deterministic adapter inference from analysis-only Reward-rerank@8. Mechanism controls show that recoverability-specific construction is clearest on FOLIO, while ProofWriter also benefits from simpler answer-filtered same-question contrastive supervision. These results support a scoped claim: proof-first models can be improved by training them to preserve recoverable problem semantics, rather than only optimizing final-answer correctness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.