acceptodds
Under review as a conference paper at ICLR 2027

Interleaving and Process-Supervising Logical Reasoning with Verifiable Reasoning Structure

Abstract

Large language models have demonstrated strong reasoning capabilities through Reinforcement Learning with Verifiable Rewards (RLVR). However, RLVR typically uses outcome correctness as a binary reward and lacks a strict check on the logical validity of intermediate steps, which matters more in logical reasoning. In this paper, we propose VERST (VErifiable Reasoning Structure Training), a training framework that obtains a process supervision signal from a symbolic solver's strict verification. The framework introduces a lightweight verifiable structure of atomic facts and explicit inferences, which can be easily translated into SMT-LIB format and verified by a symbolic solver. Based on this, our two-stage approach applies: In Reasoning Structure Learning stage, the model is trained to interleave this structure with natural language by supervised fine-tuning. In Process-supervised RLVR stage, reinforcement learning is performed with a process reward scored by a symbolic solver on the generated structure. On five logical reasoning benchmarks, VERST improves Qwen3-4B-Instruct-2507 from to , outperforming Qwen3-235B-A22B () and becoming competitive with DeepSeek-V4-Pro-0813 () and GPT-5.6-sol (). Further analysis shows that it improves process quality and brings the impact to other reasoning domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.