Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation
Abstract
A video can exhibit convincing motion and photorealism yet still fail immediately when visual text collapses. Unlike generic scene content, visual text is exceptionally unforgiving in video generation, where minor stroke corruption, temporal instability, or editing errors instantly break legibility and realism. Existing benchmarks overlook this challenge by treating text as incidental scene content or relying on static OCR metrics that ignore temporal dynamics. We introduce **VidScribe**, a unified diagnostic benchmark for visual text across four core generation regimes: writing from language (T2V), transferring text identity from reference (R2V), sustaining consistency under dynamics (I2V), and editing localized text within video (V2V). VidScribe provides 803 human-verified samples instantiated over a 12-axis conditionally orthogonal factor space spanning Intrinsic Text Properties, Physical Imaging Conditions, and Temporal Behavior. To enable reliable evaluation, we build a track-grounded, gated suite with 11 shared metrics and 2 task-specific probes that enforce strict measurability conditions. Extensive benchmarking across 11 representative commercial and open-source systems reveals that video text capability is non-monolithic, with a clear decoupling between content recognition and stroke-level glyph correctness. Performance is highly task-asymmetric, with I2V sustaining text most reliably and V2V editing emerging as the primary bottleneck. Counter-intuitively, performance degradation is more concentrated on a small subset of text-centric structural and temporal factors than on adverse imaging conditions. Further diagnostic probes show that visual reference improves glyph and typographic fidelity rather than content accuracy, while localized editing struggles to isolate target text without corrupting undeclared source text. Beyond evaluation, VidScribe also provides an actionable training signal, where benchmark-aligned preference optimization measurably improves visual text generation. We release the benchmark at https://huggingface.co/datasets/anon-submission-7k2/vidscribe-benchmark.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.