acceptodds
Under review as a conference paper at ICLR 2027

SDE Trajectory Steering with Spatial-Aware VLM Self-Verification: Test-Time Scaling for Image Editing

Abstract

Test-time inference strategies for diffusion-based editing currently face distinct bottlenecks. Standard ODE sampling is rigid and cannot correct early errors; Best-of-N scaling is computationally prohibitive; and recent tree-search alternatives introduce external reward semantic discrepancies, devolving into undirected stochastic sampling via directionless prune-and-random-restart mechanisms. To address these compounding limitations, we propose SdeSteer, a novel training-free test-time scaling framework. Our approach replaces blind noise injection with an adaptive prune-and-steer mechanism. By progressively pruning misaligned sampling paths and dynamically modulating the classifier-free guidance (CFG) scale based on an intermediate verification heuristic, SdeSteer actively steers SDE-induced noise to synthesize novel, optimized sampling trajectories. Furthermore, we achieve a zero-overhead architecture by dual-purposing the native Vision-Language Model (VLM) backbone as both the text encoder and an internal verifier. To ensure reliable evaluation without auxiliary dependencies, we introduce a spatially-aware verification strategy that effectively mitigates the self-enhancement bias typical of same-model assessment. Extensive experiments demonstrate that our actively guided test-time compute translates directly into superior generative alignment and visual quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.