acceptodds
Under review as a conference paper at ICLR 2027

Learning Programmatic Chart Editing through Privileged On-Policy Self-Distillation

Abstract

Chart editing requires a multimodal large language model (MLLM) to translate a reference chart and a natural-language instruction into executable plotting code. Existing approaches often rely on powerful proprietary models to synthesize reasoning data for supervised fine-tuning (SFT) and evaluate rollouts during group relative policy optimization (GRPO), making supervision costly to scale. To reduce this reliance, we introduce a two-stage fine-tuning framework that combines SFT on verified reasoning and code pairs with on-policy self-distillation (OPSD), where a frozen copy of the fine-tuned model receives privileged information (PI) to supervise a student model given only the original input. First, for SFT data construction, we design a reasoning trajectory generation pipeline guided by privileged information (PI). Specifically, we annotate the reference chart to obtain a structured textual description and provide it, together with the reference chart and editing instruction, to a low-cost MLLM of comparable capability to Qwen3.5-9B (the base model) to obtain the reasoning trajectory. This yields a training set of 266k verified trajectories, which, after SFT of the base model, brings improvements of over 10% on both ChartEdit and ChartMIMIC. Next, after SFT, OPSD further improves the model through dense token-level feedback, without requiring a separate larger teacher or repeated calls to an external reward evaluator during optimization. We conduct experiments to systematically study how the choice of PI and SFT initialization affect OPSD. On ChartEdit and ChartMIMIC, our model achieves relative improvements of 17% and 32%, respectively, over the base model and delivers the best overall performance among the open-source models of comparable size. Our comparisons favor precise PI aligned with the desired output, with target code outperforming both reference information and combinations containing additional context. Direct OPSD degrades performance and exhibits substantial PI leakage, whereas SFT initialization reduces the leakage rate from over 40% to approximately 2% and enables further performance gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.