–RSG: Support-Critical Reasoning Supervision from Self-Certifying Theorem Structures
Abstract
Large language models (LLMs) can often derive plausible conclusions from available premises, yet remain prone to asserting conclusions when indispensable support is missing. Existing reasoning supervision primarily emphasizes whether a conclusion follows, providing much weaker supervision about the boundary between sufficient and insufficient evidence. We investigate self-certifying theorem structures generated by the automated theorem generator as a formally grounded source of LLM reasoning supervision, and introduce -based Reasoning Supervision Generation (–RSG) to exploit their support-critical structure. Each theorem has an exact support boundary: its complete premise set entails the conclusion, whereas removing any indispensable premise makes the available support insufficient. –RSG deterministically compiles this structure into complementary supervision for conclusion inference, support sufficiency, support-based abstention, missing-support identification, and support repair, with labels justified by construction rather than external theorem verification. Experiments with Qwen2-1.5B show that self-certifying theorem supervision substantially improves controlled reasoning, while full –RSG further develops support-critical behaviour. On structurally held-out theorem configurations, –RSG achieves 98.9% inference accuracy and 100% support-sufficiency accuracy while reducing the unsupported conclusion rate from 70% for the base model to 1%. It further achieves 97% F1 for missing-support identification and 99.5% F1 for support repair. These results demonstrate that self-certifying theorem structures can provide effective, formally grounded supervision for learning both what follows from available evidence and when that evidence is insufficient to license a conclusion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.