VPBRANDLOOP: AIGC ADVERTISING IMAGE GENERATION WITH CROSS-SCENE VANISHING-POINT CONSTRAINTS
Abstract
Computational advertising requires scalable creative generation that preserves scene structure and brand identity while supporting reliable pre-deployment quality control. We present VPBrandLoop (Vanishing-Point and Brand-conditioned Closed-Loop Advertising Generation), a framework for multi-scene AIGC advertising image generation that addresses perspective drift, under-utilized visualbrand conditions, and the absence of verifiable output-side geometric control. A precomputed dominant vanishing point (VP) and perspective attributes are injected into a latent diffusion model through multi-scale geometric residuals, while product and logo references are encoded as visual brand tokens and fused through decoupled cross-attention. Training uses brand-region-weighted diffusion reconstruction, spatial-gating alignment, and auxiliary brand-discriminative supervision. At inference time, a line-segment detector (LSD) and RANSAC re-estimate the dominant VP of each generated candidate; spatial-consistency quality gating and confidence-weighted geometric feedback then drive candidate resampling. We additionally use a CLIP-based zero-shot protocol to evaluate the multimodal semantic retrievability of brands, product categories, and visually observable attributes. On 1,000 AdsGeo-10K test images evaluated across 12 methods, VPBrandLoop achieves an N-VPE of 0.2558 and an SCPR of 16.9%. Relative to ControlNet,the strongest external baseline on the corresponding metrics, N-VPE decreases by 20.3% and SCPR increases by 4.2 percentage points; both paired differences remain significant after Holm correction. The evidence therefore supports a primary advantage in preserving dominant perspective consistency between the input scene and the generated advertisement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.