acceptodds
Under review as a conference paper at ICLR 2027

Learning State Transitions Improves Downstream Policy Learning in Geometry Theorem Proving

Abstract

Geometry theorem proving can be formulated as sequential action selection over an evolving proof state. Standard policy supervised fine-tuning (\policysft) teaches a model to imitate reference actions, but does not directly supervise how those actions change the state. We study whether explicitly learning these state transitions before policy training improves downstream theorem proving. We build a stateful agent environment around the open-source GenesisGeo symbolic engine. Within this environment, we first apply world-model supervised fine-tuning (\wmsft) to predict auxiliary line construction feedback, whose newly derived geometric relations define the action-induced state delta. We then perform on actions from engine-generated trajectories using the same model. Across three runs, followed by achieves 72.9% mean proof success, compared with 65.9% for alone. Additionally, after subsequent reinforcement learning, followed by and RL reaches up to 76.8%, compared with 71.2% for followed by RL. is used only during training, all policies follow the same action-selection procedure at inference, with no additional prediction stage, model call, or planning. These results show that supervised learning of action-induced state transitions can provide an effective initialization for training mathematical-agent policies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.