DesignOrbit
Abstract
Early concepts for buildings, pavilions, and public installations often exist as a single sketch or image, yet judging them requires seeing their unseen sides and how they sit in their surroundings. We study generative design in context: from one image and a text description, we sample alternative designs as loop-closed orbit videos, each proposing plausible unseen sides while keeping the design and its surroundings consistent across the full orbit. We introduce DesignOrbit, which learns this behavior without any real architectural orbit videos. Procedural cuboid buildings rendered along known camera paths, combined with loop-closed background inpainting, provide initial supervision. A structure-from-motion verifier then selects the model's own generations for self-training based on frame registration and camera coverage, and later supplies preference pairs for direct preference optimization, extending training beyond the procedural shapes while increasing camera coverage. On architecture, camera coverage rises from to , against for the strongest baseline, while diversity in unseen rear views is retained. Although trained only on synthetic building videos, the model also generates orbits for sketches, playgrounds, ships, interiors, and urban scenes without further fine-tuning. In a study with 50 participants, 80% with a background in architecture, design, or 3D visualization and over half with five or more years of experience, DesignOrbit was preferred over the baselines in 91% of ratings that expressed a preference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.