STIR-FM: Segment-wise Trajectory Interpolation Refinement for Flow Matching
Abstract
Flow-matching models require multiple sequential model evaluations for text-to-image generation. We propose STIR-FM, a training-free post-sampling framework that refines a cached few-step trajectory by inserting model evaluations into intervals selected using velocity variation and propagating local velocity corrections. STIR-FM scores each interval by the velocity change across it, inserts an evaluation at the midpoint of the highest-scoring interval, and corrects the cached velocities downstream from a coordinate-wise spatial slope estimated from three evaluated reference pairs. An initial trajectory costing model evaluations and inserted evaluations has a total cost of NFEs, with only new evaluations needed when the initial trajectory is already available. Fixed-budget experiments on FLUX.1-dev and Stable Diffusion 3/3.5 show improved mean ImageReward over the initial trajectories; at matched total NFE, the refined means approach or exceed direct sampling once the initial trajectory is moderately resolved. A separate STORK experiment refines only samples falling below their per-sample STORK-12 CLIP reference and selects the highest-CLIP output in each history, reducing counted generator NFEs by % while mean held-out ImageReward remains numerically close to that of fixed STORK-12.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.