Learning State Transitions Improves Downstream Policy Learning in Geometry Theorem Proving
Abstract
Geometry theorem proving can be formulated as sequential action selection over an evolving proof state. Standard policy supervised fine-tuning (\policysft) teaches a model to imitate reference actions, but does not directly supervise how those actions change the state. We study whether explicitly learning these state transitions before policy training improves downstream theorem proving. We build a stateful agent environment around the open-source GenesisGeo symbolic engine. Within this environment, we first apply world-model supervised fine-tuning (\wmsft) to predict auxiliary line construction feedback, whose newly derived geometric relations define the action-induced state delta. We then perform on actions from engine-generated trajectories using the same model. Across three runs, followed by achieves 72.9% mean proof success, compared with 65.9% for alone. Additionally, after subsequent reinforcement learning, followed by and RL reaches up to 76.8%, compared with 71.2% for followed by RL. is used only during training, all policies follow the same action-selection procedure at inference, with no additional prediction stage, model call, or planning. These results show that supervised learning of action-induced state transitions can provide an effective initialization for training mathematical-agent policies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.