PhyScale: Physically Grounded Object Size Correction in Text-to-Image Generation
Abstract
Generated images can appear photorealistic while depicting objects at implausible relative physical sizes. Correcting these errors requires distinguishing physical size from its depth- and pose-dependent appearance. We introduce PhyScale, a framework that formulates size correction as geometric inference and planning, separating physical-size changes from their image-space realization. Using typical real-world object dimensions from a knowledge base and estimated scene geometry, it computes size, depth, and placement commands for an image editor controlled through spatial rotary positional embeddings. To learn the associated appearance changes, we adapt the editor on only 2,468 paired renders from controlled 3D object transformations. The resulting editor achieves the best target-fidelity and identity-preservation metrics among the compared editors on prescribed edits. Combined with our geometric inference and planning front end, it delivers aggregate GenScale correction performance comparable to a state-of-the-art correction pipeline at approximately one-twentieth of its estimated inference cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.