acceptodds
Under review as a conference paper at ICLR 2027

CACO: Cost-Aware Dynamic Caching with On-Policy Distillation for Accelerating Diffusion Models

Abstract

Diffusion Transformers (DiTs) achieve strong performance in high-resolution visual generation, but their iterative denoising requires repeated backbone computation, resulting in substantial inference cost. Feature caching accelerates sampling by reusing intermediate features and skipping backbone computation at selected steps. However, their practical effectiveness is limited by scheduling overhead and the accuracy of approximated features, with existing feature predictors further suffering from train–test mismatch caused by teacher-forcing training. In this paper, we propose CACO, a global–local framework that integrates global cache scheduling with local feature prediction to improve both scheduling efficiency and feature estimation accuracy. Specifically, we propose a Cost-Aware Global Scheduler (CAGS) that first allocates a fixed budget of refresh anchors and dynamically relocates refresh positions through sparse probing while preserving the refresh budget, thereby reducing the overhead of dynamic scheduling. For feature estimation, we present Time-Conditioned Feature Predictor (TCFP), which is first trained with supervised fine-tuning (SFT) on full-computation reference trajectories and subsequently refined with On-Policy Distillation (OPD) on cached trajectories to mitigate train–test mismatch during inference. Extensive experiments on image and video generation demonstrate that CACO substantially reduces inference cost while maintaining generation quality, enabling efficient practical deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.