Skeleton-Preserving Coefficient Refinement for Multimodal Physical Formula Discovery
Abstract
Discovering physical formulas from multimodal scientific data has recently been advanced by vision-language models (VLMs) that map images, text, and trajectories to symbolic expressions. However, the prevailing post-hoc coefficient-fitting strategy (SR ) systematically corrupts the predicted formula skeleton: in our analysis, SR alters the skeleton in of strictly comparable samples and decreases the structural score in of altered cases (mean drop ), reducing average from to . This reveals a fundamental structure-precision dilemma: refinement procedures that jointly search over structures and coefficients can, and in practice do, trade skeleton correctness for marginal coefficient gains. We address this by reformulating refinement as a constrained problem in which the upstream skeleton is fixed and only coefficients are updated, making structural integrity an architectural invariant rather than an optimization outcome. We propose HAD (Hybrid Autoregressive-Diffusion), a skeleton-preserving coefficient refinement module that reuses an upstream VLM's symbolic token sequence and refines only the continuous coefficients at numeric positions via an autoregressive-conditioned diffusion head. A decoupled two-stage training schedule resolves the AR–diffusion parameter mismatch. On PhysSymbol benchmark, HAD improves from to and from to over the strongest VLM baseline, while eliminating SR 's structural corruption and reducing refinement cost from to per sample.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.