AdaShare: Adaptive Cross-Frame Sharing for Training-Free Consistent Image Generation
Abstract
Consistent visual story generation requires maintaining subject identity and visual style across images under different scene prompts. Existing training-free methods improve cross-frame consistency by sharing reference features across frames. However, unreliable feature sharing can lead to incomplete subject generation and degraded image quality. To address these issues, we propose a training-free framework for adaptive cross-frame feature sharing in consistent text-to-image generation, termed AdaShare. To improve subject completeness and consistency, AdaShare extracts reliable subject masks from cross-attention maps and activates subject feature sharing after subject regions stabilize. We further introduce Adaptive-Capacity Optimal Transport Attention (ACOT Attention), which formulates Extended Attention as an optimal transport problem. By imposing adaptive capacity constraints derived from the areas of the subject masks, ACOT Attention prevents excessive attention concentration on reference tokens caused by token imbalance, enabling more effective utilization of cross-frame features. Furthermore, to maintain style consistency, we introduce Covariance Value Alignment, which applies whitening and coloring transformation to the low-frequency components of value representations while leaving high-frequency structures unchanged. Extensive experiments on ConsiStory+ demonstrate that AdaShare achieves superior performance over training-free baselines in core consistency metrics and human evaluations, while maintaining prompt fidelity and image quality. The source code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.