CPUNeSy: Controlling Model Writes for Reliable Neuro-Symbolic Reasoning
Abstract
Large language models are highly effective at recalling statistical patterns, but their accuracy degrades sharply when answers must be derived rather than recalled, especially on deep multi-hop chains. Delegating derivation to deterministic symbolic executors shifts the reliability question to whether model-generated premises are supported by the source. We introduce CPUNeSy (Certified Predicate-Unit Neuro-Symbolic Networks), a serving architecture that controls model writes to symbolic state through a task-defined predicate interface and a certificate gate, and abstains when grounding passes disagree. A component analysis separates deterministic execution from restricted grounding, agreement, and source rechecking. Our experiments show that deterministic execution accounts for most of the accuracy recovery on derivation-heavy tasks, whereas controlled writes primarily improve selective reliability by withholding unsupported or inconsistent answers at the cost of coverage. On controlled multi-hop stress tests in law and formal mathematics, deterministic execution recovers most of the accuracy gap over chain-of-thought and retrieval baselines, with full-pool gains up to 35.0 points. Certification behaves as a selective-serving control rather than an accuracy mechanism: with grounding traces held fixed on ContractNLI, source rechecking removes a quarter of DeepSeek's wrong answers that survive two-vote agreement, at a measurable coverage cost. When abstention is expensive, routing withheld cases to an explicitly uncertified same-model fallback raises full-pool accuracy on MedCalc-Bench Verified by and points for Seed and DeepSeek; these gains do not belong to the certified serving channel. On LeanDojo Benchmark, kernel-restricted pools match BM25 recall@ (). Gains depend on the grounding model's error regime: bias-dominated grounders benefit less, consistent with the voting bound we state formally. The resulting certificates guarantee derivational validity relative to admitted premises; semantic faithfulness to natural-language sources remains conditional on the source checker, and prospective validation of the deployment protocol is future work.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.