StreamAccel: Automated Inference Acceleration for Streaming Causal Video Models
Abstract
Interactive video generation requires fast, sustained causal inference to provide timely feedback to user inputs. Yet accelerating a new model still requires substantial manual tuning: practitioners must not only select suitable acceleration techniques but also determine their order and combination. Because earlier optimizations can alter the execution conditions of later techniques, choices that perform best in isolation may not be optimal when combined. Moreover, differences in model structure, streaming state, and deployment make existing configurations difficult to transfer directly, requiring these decisions to be revisited for new models. To address these challenges, we introduce StreamAccel, an automated inference optimization system for streaming causal video models. Its staged workflow defines the optimization tasks and prerequisites for each stage, guiding an agent to select, compose, and validate acceleration techniques. Steady-state measurements assess the benefits of these combinations for diffusion-transformer (DiT) generation and VAE decoding. Complementing this workflow, a reusable technique knowledge library captures applicability conditions, composition experience, and failure modes from prior tasks, guiding optimization on new models and enabling engineering experience to be reused across models. Experiments demonstrate that StreamAccel consistently reduces steady-state latency for both generation and decoding across different causal video models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.