Shuqi: Support-Aware Executable Reasoning for Representation–Execution Diagnosis
Abstract
Final-answer accuracy reveals whether a reasoning system is wrong, but not whether the error arises from constructing a usable representation or from executing semantics that are already explicit. We study this distinction with a controlled Raw-NL → Same-IR → deterministic diagnostic interface and semantic-stability audits that separately expose representation-associated recovery, residual execution-side error, and instability under fixed semantic content. These diagnostics reveal substantial effects: residual execution gaps reach 23.0 percentage points, while lossless serializations of identical canonical representations disagree on up to 30.8% of disjoint-source examples. Based on these findings, we propose Shuqi, a task-grounded support-aware executable reasoning system that compiles supported inputs into executable representations, performs deterministic task-defined execution with replay-checked evidence, and releases the symbolic result only when a task-specific support contract succeeds; otherwise, it falls back to a frozen generative reasoner. On ProofWriter, Shuqi preserves all 41 symbolic fixes while introducing no regressions, reaching 98.0% accuracy, and achieving 100% on both ProntoQA and LogicalDeduction. These results show that reliable reasoning requires controlling not only what semantic representation is constructed, but also how that representation is executed and when an executable result should be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.