Reallocate, Rescale, Retain: Test-Time Latent Control for Image Generation
Abstract
Optimizing continuous text and image latents at test time can improve image generation in unified multimodal large language models without updating pretrained parameters. Yet these latent states evolve during optimization, while the policy controlling their refinement often remains fixed. This mismatch can waste updates on locally saturated positions, ignore differences in update reliability, and discard better intermediate images by returning the final iterate. We introduce TRILO (Triple-Control Latent Optimization), a state-dependent method that controls where to refine, how to scale update signals, and which evaluated image to return. Reallocate uses predictive entropy to shift a fixed-size text-latent refinement window away from local saturation. Rescale adjusts update signals within the selected window, applying more conservative scaling at less confident positions. Retain returns the highest-reward evaluated image if the stopping criterion remains unmet after budget exhaustion, without additional generation or reward evaluation. Together, these controls adapt how a fixed optimization budget is deployed rather than increasing it. Local theoretical analysis and controlled diagnostics support these design principles. Experiments on GenEval, T2I-CompBench, and WISE demonstrate consistent improvements in compositional and knowledge-intensive generation on Janus-Pro-1B and Janus-Pro-7B. With benchmark evaluators as online rewards and matched maximum update budgets, TRILO achieves higher scores and 12.8–24.0% lower measured end-to-end runtime than a fixed-policy baseline across all three benchmarks on the 7B model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.