Does Memorability Guidance Transfer Beyond the Guiding Critic? Evidence from Latent Diffusion
Abstract
Image memorability is the likelihood that an image will be remembered after viewing. Neural memorability predictors estimate this likelihood and can guide diffusion models toward images with higher predicted memorability. We call the predictor that provides this guidance the guiding critic. Prior work shows that such guidance can also raise average scores from predictors not used for guidance. What remains unclear is (i) whether improvements in predicted memorability transfer consistently beyond the guiding critic at the individual-image level, (ii) whether these improvements arise from the intended guidance direction rather than generic changes introduced by generation or reconstruction, and (iii) whether gradient-based guidance justifies its added computation compared with simple sampling alternatives. We investigate these questions using training-free universal guidance in latent diffusion for both text-to-image generation and real-image editing. The critic supplies gradients during sampling, while the other predictors serve as held-out evaluators that only score the final images. With AMNet as the critic, guidance raises the mean ResMem, ViTMem, and PerceptCLIP scores by 2.87%, 1.69%, and 1.67% on 200 COCO prompts fixed before generation, while reducing CLIP prompt similarity by 1.88%. These gains hold across four seeds on DrawBench and COCO, and on DrawBench they hold whichever of AMNet, ResMem, or ViTMem acts as the critic. At the image level, however, the held-out predictors often disagree, with all three scores rising together on only 69 of 200 COCO image pairs. To separate the effect of guidance from the changes that noising and denoising make on their own, we edit real SUN photographs and compare each edit with a matched unguided reconstruction. Forward guidance raises and reverse guidance lowers the held-out scores, while a random push of the same size has near-zero effect, so the change follows the direction of the critic's gradient. Generating four unguided images and keeping the one the critic scores highest is competitive with guidance on the held-out scores, keeps prompt similarity higher, and runs faster in our benchmark. A five-prompt diagnostic on Stable Diffusion XL shows that predictor scores can rise even when images are visibly malformed. Together, these results show that memorability guidance produces real but modest gains beyond its critic, which hold on average but not reliably per image, cost prompt fidelity, and offer little over selecting the best of a few samples.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.