CodaCache: Coordinated Adaptive Branch Caching for Efficient Image-to-3D Generation
Abstract
Image-conditioned diffusion and flow models generate detailed 3D assets from a single image. Full classifier-free guidance (CFG) computes conditional and unconditional predictions at every denoising step, increasing generation cost. The effect of reusing one prediction depends on whether the other is recomputed or reused. In our Hunyuan3D 2.1 diagnostic, reusing both predictions yields a smaller deviation from Full CFG than reusing the conditional prediction while refreshing the unconditional one. At a fixed state and timestep, reuse errors with similar directions and magnitudes can partly offset each other in the CFG prediction. This paper proposes CodaCache, a training-free causal controller that adaptively coordinates CFG branches. It estimates the effects of omitting the unconditional prediction and reusing the conditional one from completed denoising history, using latent motion and time separation. It combines these estimates to choose Full CFG, conditional-only refresh, or conditional-only reuse before current branch evaluation, without extra denoiser calls. On TRELLIS.2 and Hunyuan3D 2.1 with Toys4K and an Objaverse-XL subset, CodaCache uses 37.6-44.7% of Full-CFG denoising FLOPs and achieves 2.34-2.73 end-to-end generation speedup, excluding offline calibration. Across valid outputs, mean Chamfer distance increases by at most 0.15% and mean F-score decreases by at most 0.10 percentage points relative to Full CFG.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.