Diverse Narratives, Preserved Content: Controllable Affective Image Generation
Abstract
Controllable affective image generation aims to generate images from content descriptions and target emotions. A given content description and emotion target can support multiple visual narratives, but such diversity should arise from variations in actions, interactions, and situations rather than changes to protected content. However, existing methods struggle to distinguish content that must be preserved from narrative elements that can vary. To address this challenge, we propose a content–narrative decomposition framework that preserves specified entities, attributes, and required relations while enabling diverse narrative expressions. Specifically, scene-graph supervision and cross-reconstruction encourage composable content and narrative representations. Conditional flow matching models a distribution of latent narrative prompts, which, together with protected content and target valence–arousal values, guide an emotion injection module to update narrative features. These features are then recombined with the content representation to condition a frozen SDXL generator. To support one-to-many narrative learning, we augment the existing training data with diverse emotional descriptions for each content–emotion pair. Experiments demonstrate that our method reduces content violations and improves constraint-satisfying narrative diversity while maintaining emotion controllability. Ablation studies confirm the complementary roles of content–narrative decomposition and flow-based narrative modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.