AO-FLOW: LEARNING TO USE AND OPTIMIZE GENERATED VIEWS FOR 3D ASSETS
Abstract
Orbit videos provide candidate observations for single-image 3D generation, yet different regions need different views and adjacent frames can repeat the same error. AO-Flow first adapts a 3D receiver to use the full 49-frame video: sparse- structure sampling retains region-specific source identities and attention masses, and Shape rereads those frames with its own queries and a soft spatial prior. Se- quential supervision trains the source producer before freezing it for Shape adapta- tion. AO-Flow then fixes the receiver and uses final-asset feedback to update only video attention LoRA. The regional interface reaches H-F 0.639 versus 0.613 for a global four-source reader at similar receiver cost; full 49-frame reading attains slightly higher H-F at greater cost. Asset feedback raises H-F to 0.669 while trad- ing some input-view agreement. Inference uses one image, one video and one unselected asset
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.