acceptodds
Under review as a conference paper at ICLR 2027

Faster3D: Exploring the Roles of Latent Tokens for Faster Image-to-3D Generation

Abstract

Flow-based image-to-3D models repeatedly predict velocities for all latent tokens during sampling, although the need for recomputation can vary across tokens. Image-to-3D generation is primarily concerned with the sparsity problem that the surrounding space of the formed object always remains empty. Our analysis of TRELLIS sparse-structure generation shows that non-surface tokens associated with empty regions generally have smaller velocity magnitudes and directional changes than surface tokens associated with occupied regions. These observations motivate selective velocity reutilization. However, omitting selected tokens from self-attention can degrade geometry because their representations still provide context for other tokens. In this paper, we propose Faster3D, a training-free acceleration framework with two components. The Temporal Token Decoupling which instructs the reutilization of velocity predictions, and the Cached Context Restoration retains cached keys/values at their original positions for sampling. These two components reduce redundant velocity computation while preserving the attention context provided by selected tokens. Experiments on the Toys4K dataset show that Faster3D consistently accelerates diverse image-to-3D backbones by 1.30×–1.95× speedups in the targeted structure or geometry sampling stage, with geometry and appearance metrics remaining close to the full-compute baselines. These results highlight context-preserving token-wise velocity reutilization as an effective technique for image-to-3D accelerations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.