acceptodds
Under review as a conference paper at ICLR 2027

Struct-VidOCR:Instance-CentricOCREvidenceforVideoCaptioningviaRL Alignment

Abstract

While Vision Language Models (VLMs) excel at static image OCR, Video OCR remains difficult because text moves, is occluded, and updates over time, disrupting persistent tracking and causing temporal hallucinations. We propose Struct-VidOCR with Structured OCR Segments, explicit lifecyclerecordsthatpreservetranscription,natural-language position, sampled temporal span, semantic and source attributes, motion, occlusion, persistence, and content variants as reusable OCR evidence for a self-contained caption. We align this record–caption output using Group Relative Policy Optimization (GRPO) and a deterministic matching-derived reward over OCR fidelity, temporal and attribute grounding,andcaptionconsistency.Trained on only1.5k automatically quality-controlled pseudo-labeled clips, our 8B model improves Qwen3-VL-8B by 2.5 points on MME-VideoOCR and 3.7 points on the caption-compatible VidText subset, while achieving performance comparable to Qwen3-VL-235B-A22B-Instruct. These results user igorous Caption-Injected QA: a fixed text-only reader answers benchmark questions from the generated caption without video frames; the predicted record provides additional downstream utility. Thus, Struct-VidOCR validates a data-efficient, structurally aligned video-to-text evidence interfaceunder this controlled readout.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.