Plan What the Reply Needs: Requirement-Guided Synthetic Supervision
Abstract
High-quality synthetic supervision requires more than correct content: a response must also decide what to ask, explain, qualify, or correct for a particular user. We study whether making these response requirements explicit during data construction yields improvements that students can learn. We present a plan-then-revise approach built around ASMF, a planning interface defined without domain terms that records concrete requirements, relevant conversational states, response actions, and their expression. A planner reads only the conversation, and a reviser uses its plan to improve the teacher's draft. Construction uses neither evaluation rubrics nor candidate selection, and students learn only from the resulting conversation-response pairs. We test this on HealthBench, a medical benchmark whose physician-written rubrics grade what each reply needs. On 4,000 training conversations, plan-guided revision raises the teacher's rubric score by 6.8 percentage points (95% CI [6.2, 7.5]) over generic revision, which applies the same revision rules without a plan. Across three student backbones from two model families (Qwen3.5 4B and 9B, Ministral-3 8B), students trained on plan-guided revisions exceed those trained on generic revisions by 5.0 to 7.1 percentage points in all six comparisons on held-out HealthBench Hard and the independently released HealthBench Professional, with every 95% confidence interval excluding zero, while generic revision gives no reliable improvement over direct answers. The gains are largest in completeness and context awareness. These findings support explicit response planning as a practical way to construct synthetic supervision, with benefits that persist when students answer without any plan.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.