Behavior-Aligned Trust Regions for SVG Reinforcement Learning
Abstract
Reinforcement learning from verifiable rewards (RLVR) for SVG generation typically constrains policy updates in token space. Yet token-space constraints can be poorly aligned with visual behavior: the same drawing may admit many syntactically distinct programs, while a small textual edit can substantially alter the rendering. This mismatch creates opportunities for reward hacking, where the policy exploits shortcuts that preserve a proxy reward while degrading the drawing. We study whether a behavior-aligned representation space can complement tokenlevel constraints in SVG RLVR, and instantiate this space with a variational latent model built on a frozen SVG-finetuned LLM encoder. The Gaussian posterior provides a probabilistic local representation with closed-form divergences and uncer- tainty estimates, while a lightweight metric adapter calibrates the local geometry using programmatic SVG transformations. We first show that frozen encoder hidden states do not reliably track rendering similarity, motivating the learned metric. We then combine the calibrated behavior representation with a token-KL syntax anchor in a hybrid constraint, using an uncertainty-adaptive local radius. On SVG RLVR, the resulting method reduces the two studied collapse behaviors (viewBox and length collapse) while preserving or improving generation quality, and the representation remains useful on unseen sources and out-of-domain SVGs, while performance is lower on held-out transformation families.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.