ViNAYA: Yoked Schedules for Streaming Non-Autoregressive Captioning
Abstract
A streaming captioner must decide how much of a clip it has seen before writing every word. Autoregressive models answer this by construction at the cost of one sequential pass per token. Non-autoregressive models are faster but cannot stream because no prefix is guaranteed final until the last pass and nothing can leave early. We introduce ViNAYA (Vision Non-Autoregressive Yoked Architecture), which closes this gap by turning the write order into a declared schedule rather than a learned behaviour. A monotone text frontier commits and emits spans from left to right, while a yoked monotone source horizon reveals the clip at a matching pace. Both frontiers are closed mathematical forms fixed before decoding begins, making them provably safe under arbitrary arrival. Therefore, no emitted token can depend on source that has not yet arrived, and every emitted span is final the moment it leaves. ViNAYA's two-knob design permits separate control over both the source reveal and the text schedule, with the two knobs joined by a ratio grounded in corpus geometry. This coupling yields a predictable exposure budget of , which we validate against measured exposure on three video captioning corpora with three frozen encoders. On the same corpora, ViNAYA reduces Average Lagging by 43 to 50 percent against the strongest NAR streaming baseline, at quality statistically indistinguishable in most cases and with core architectural elements shared by construction. While an autoregressive wait- baseline retains a small quality edge on two corpora, ViNAYA compensates that with up to 2.7 times fewer forward passes per window and a sequential cost that stays bounded as captions grow. The per-token structure of autoregression forces its inference time to grow with output length, and on ActivityNet, where captions are longest, the autoregressive baseline requires roughly 408 seconds of test-split inference against 188 seconds for ViNAYA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.