ATHENA: Test-Time Steering for Count Fidelity in Text-to-Image Diffusion Models
Abstract
Text-to-image diffusion models achieve high visual fidelity yet exhibit systematic failures when prompts specify explicit object counts. We introduce ATHENA, a model-agnostic, training-free framework for improving count fidelity through test-time steering. ATHENA constructs a cardinality-specific contrastive steering signal by varying the numerical condition in the input prompt, enabling targeted count intervention during denoising while preserving the underlying semantic content. However, different generation trajectories develop different count errors during denoising, making a fixed steering signal insufficient. To address this, we further incorporate intermediate count feedback to identify emerging cardinality errors, construct trajectory-specific corrections, and adjust the steering strength based on the observed response. Across multiple text-to-image diffusion backbones and counting benchmarks, we show that ATHENA consistently improves count fidelity over unsteered generation and existing counting methods, with particularly strong gains in challenging counting settings while maintaining favorable accuracy-runtime trade-offs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.