FLUID-Omni: Parallelizing Streaming Omni Generation with Commit-Repair-Write Diffusion
Abstract
Autoregressive omni-modal models naturally support streaming generation because each newly produced token extends a stable causal prefix that can be immediately consumed. However, their sequential decoding process limits generation efficiency. Discrete diffusion decoding provides a promising alternative through parallel generation, yet directly applying it to omni-modal generation introduces a fundamental challenge: future predictions may become available before they are reliable enough for irreversible commitment. We identify this prediction–commitment mismatch as a key barrier to efficient parallel streaming generation. We introduce FLUID-Omni, a causal diffusion framework that bridges parallel generation and streaming reliability through a Commit–Repair–Write paradigm. Instead of treating all predicted future tokens as equally reliable, FLUID-Omni selectively commits stable predictions, locally repairs inconsistent regions while preserving reusable context, and continuously writes new predictions in parallel. To align training with streaming inference, we develop a multi-stage adaptation strategy that equips the model with causal parallel prediction and pipeline-consistent refinement capabilities. Furthermore, we introduce prior-calibrated codec decoding to mitigate biased confidence estimation in audio generation. Experiments across semantic, acoustic, and streaming evaluations demonstrate that FLUID-Omni substantially improves generation efficiency while preserving omni-modal generation quality. These results show that parallel diffusion can achieve reliable streaming generation when prediction and commitment are explicitly decoupled.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.