Generative World Renderer at the Speed of Play
Abstract
G-buffer-conditioned generative rendering synthesizes RGB frames from geometric and material buffers exported by a game engine, with appearance controlled through text prompts. By retaining engine-side geometry, physics, and gameplay logic, this approach provides a foundation for interactive, user-controllable visual worlds. However, the computational cost and fixed-window formulation of existing diffusion-based renderers limit their use in real-time applications. We introduce StreamRenderer, a real-time generative renderer that accelerates a 50-step diffusion-based teacher from 0.56 FPS to 31.54 FPS on a single NVIDIA H200 GPU. StreamRenderer combines autoregressive streaming, progressive four-step distillation, and lightweight distilled codecs for efficient latent encoding and frame reconstruction. It retains the teacher’s G-buffer and text-prompt interfaces while supporting continuous rendering and online prompt updates over input streams of arbitrary length. We evaluate StreamRenderer in terms of content preservation, temporal consistency, cross-window stability, prompt controllability, and runtime efficiency. The results demonstrate substantial computational savings while largely preserving content fidelity and prompt controllability, with a tradeoff in temporal stability relative to the bidirectional teacher. Integrated with a game engine, StreamRenderer enables a fully playable, prompt-controllable generative rendering system running at 30 FPS, bringing generative world rendering to the speed of play.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.