acceptodds
Under review as a conference paper at ICLR 2027

VeriPO: Closing the Reasoning-Supervision Gap in Document Reinforcement Learning

Abstract

Reinforcement learning with rule-based, verifiable rewards has become the dominant post-training recipe for document vision-language models, yet it treats the decoded trajectory as an undifferentiated scalar-reward target. We observe that this practice clashes with the internal structure of document-QA outputs: a trajectory decomposes into parse, evidence, reason, and answer segments of sharply different verifiability, and the reasoning segment admits no rubric at all. Rule-only objectives therefore leave reasoning tokens with no within-trajectory discriminative signal, producing slow, variance-dominated training, a phenomenon we call the reasoning-supervision gap. We close this gap with Verifiability-aware Policy Optimization (VeriPO), an on-policy reinforcement-learning algorithm whose token-level advantage combines per-segment verifiable rewards with a token-level reverse-KL signal from a frozen vision-language teacher. A rubric-gated trust coefficient, computed by scoring the teacher against the same rubric the student faces, attenuates teacher supervision wherever the teacher itself fails verification. The design admits a provable verification-safe guarantee: under an adversarially uninformative teacher, VeriPO's gradient collapses onto the rule-only baseline and never underperforms it. VeriPO further trains with a single on-policy rollout per update, eliminating the group-rollout overhead of GRPO-style training. Applied to a unified document-QA decoder that emits parse, evidence, rationale, and answer in a single pass with self-citing Document Units, VeriPO advances the state of the art on multi-page document QA, chart and infographic QA, and document parsing, at a fraction of the post-training compute of group-rollout baselines. Three diagnostic studies confirm each of its mechanisms: a component ablation, an adversarial-teacher study, and a per-segment analysis that isolates reasoning as the primary beneficiary of teacher supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.