PaRa-OPD: Pathology Reward-Aware On-Policy Distillation via Token-Level Supervision Reweighting
Abstract
Pathology vision-language reasoning requires models to connect fine-grained histologic evidence with specialized diagnostic concepts, while the task-critical information in a generated rationale is often concentrated in only a small fraction of tokens. On-policy distillation (OPD) provides an effective way to transfer reasoning behavior from a teacher on student-generated trajectories, but its standard token-level objective supervises visually grounded diagnostic tokens in much the same way as generic reasoning and connective text. This uniform treatment limits the effectiveness of OPD on pathology tasks. Existing online visual-gain weighting methods attempt to identify visually dependent tokens during OPD, but their signal can weaken when the sampled student token is already visually incorrect. To make OPD more effective for pathology reasoning, we introduce PaRa-OPD, a pathology reward-aware OPD framework for token-level supervision reweighting that directs distillation toward tokens that either depend on visual evidence or encode pathology-specific diagnostic information. PaRa-OPD trains a Pathology Reward Model (PRM) to identify these complementary token patterns in generated trajectories. The resulting PRM scores are converted into token-level weights that dynamically reweight the OPD objective, strengthening supervision where visual grounding and diagnostic knowledge are most consequential. On PathMMU, PaRa-OPD improves Qwen3.5-2B from 63.88 to 71.44 and Qwen3.5-4B from 66.96 to 74.96 over standard OPD, demonstrating strong and consistent effectiveness across model scales. Code and processed data will be released in the future.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.