Where and When: Structure-Grounded Credit Assignment for Cross-Modal Reinforcement Learning
Abstract
Reinforcement learning for visual code generation is cross-modal: the reward comes from rendered pixels, but the actions are code tokens. Most existing methods assign the same advantage to every token in a response and the same credit to every edit in an episode. Finer-grained token-level methods ignore which screen region a token affects. Yet each rendered node, such as a DOM element, links a screen region to a code span. We propose Structure-Grounded Credit Assignment (SGCA), which uses these nodes to locate where in a response and when in an episode the reward is earned. Spatial credit reweights the advantage of each node's tokens according to that node's rendering error: below-average responses are penalized most where they differ from the target, and above-average ones are reinforced most where they match it. Temporal credit scores each edit by its subsequent return minus a learned baseline for edits to that node. On Design2Code, the spatial credit outperforms all baselines, improving CLIP score by 1.53 points, and its gain grows fivefold from the simplest to the most complex pages. In multi-turn repair, the gains from the two credits stack: together they improve CLIP score by 1.37 points and reduce LPIPS by 18% on web pages, and raise the official ChartMimic score, a metric not used in any reward, by 4.3 points. All twelve pre-specified comparisons remain significant after Holm correction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.