Same Object, Different Emotions: Object-Centric Emotional Content Generation
Abstract
Text-to-image models can generate semantically coherent images, yet often struggle to make a specified object convey a desired emotion. Existing methods typically rely on scene-level cues or distort the object to strengthen emotional expression, compromising object identity and realism. We address object-based emotion-controllable image generation, where both the target object and the overall image should express the intended emotion while remaining visually plausible. To support this task, we construct Obj-EmoSet, which augments affective images with object categories, basic content descriptions, and emotion-aware descriptions. We then propose an Emotion-Aware Visual Planner that translates abstract emotions into object-compatible states, actions, and scene conditions. A Visual Realism Refiner improves composition, spatial relationships, and textures while retaining emotional cues, and an Emotion Calibration Adapter aligns emotional semantics with visual features. Together, these components enable object-level emotional expression while preserving semantic consistency and visual realism. Extensive qualitative and quantitative experiments demonstrate strong performance in object-level emotion control, semantic consistency, and image realism, advancing emotion-aware generation from coarse global modulation toward fine-grained object-centric control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.