RAVEN: Reasoning Across Premise Orders with Source-to-Response Certification
Abstract
Logical consequence is invariant to premise order, yet language models are not: reordering the same facts can change both the predicted answer and the symbolic program executed by a solver. Consequently, a locally valid proof may still certify a mistranslated input. We introduce RAVEN, a neuro-symbolic reasoner that turns premise-order sensitivity into testable evidence. It produces direct labels and formal results from source-equivalent premise orderings, validates each formalization, repairs syntax or type failures, executes each theory, and checks proof and source support. When the orderings disagree, a fixed decision rule uses their checked evidence to select a label; a separate reasoning-release certifier subsequently binds the selected formalization, proof, displayed reasoning, and answer to the source. With Qwen2.5-32B, RAVEN's accuracy-first controller reaches 88.76% macro-average accuracy across four benchmarks. Full reversal improves the fixed two-ordering study by 2.60 points over its same-order control. Under a common source-bound certification protocol on the same 600 ProofWriter test items, 46.33% (278/600) of RAVEN-Core outputs and 9.83% (59/600) of SymbCoT outputs are accepted, with no observed error in either accepted subset. Finally, on 19,790 audited finite-Horn items under registered translation evidence, the Core-attached certifier authorizes 10,901 responses with no observed accepted error and a 0.035% item-level 95% Wilson upper bound.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.