acceptodds
Under review as a conference paper at ICLR 2027

GroundForm: Learning to Ground Autoformalization with Verified Hindsight Credit and Growing Memory

Abstract

Autoformalization requires a model to reconcile natural-language intent, visual evidence, and domain knowledge in a single precise statement. We present Ground- Form, an agent that formalizes both textual mathematics and multimodal physics and learns which knowledge to retrieve along the way. It is trained with Verified Hindsight Credit (VHC): the formalizations that pass verification within a rollout group serve as targets, and a frozen reference model scores how much closer each tool interaction brings the agent to them, so that intermediate retrieval decisions receive their own credit. In addition, GroundForm maintains two verified memo- ries that grow during training. Mathematical experience can be retrieved in both domains, and physical experience supplies laws and conventions missing from the static libraries. On eight benchmarks, GroundForm reaches 80.8% average semantic accuracy in mathematics and 58.2% in physics, 57.6 and 43.6 points above its untrained Qwen3.5-9B backbone. Interventions on tools, memory, and images help locate the source of these gains. VHC increases training time per step by 16.8%, yet it needs 41% fewer GPU-hours to reach the final PhyX-AF accuracy of outcome-only training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.