Shape Meets Surface: Compositional Sketch-Text Retrieval of Textured 3D Objects
Abstract
Sketches provide an intuitive way to express object geometry, but they are insufficient for specifying textured appearance such as material, color, pattern, and surface style. Text offers complementary appearance descriptions, yet it is often less precise in constraining fine-grained geometry. In this paper, we study a new task: sketch-text guided textured 3D model retrieval, where the goal is to retrieve a textured 3D model that simultaneously satisfies sketch-based geometric intent and text-based appearance intent. To support this task, we construct ST-Tex3D, a tri-modal benchmark containing 1,005 textured chair models and 555 textured lamp models, with aligned SVG sketches, ShapeNetCore meshes, texture variants, multi-view renderings, and retrieval-oriented textual descriptions. We further propose a unified retrieval framework that jointly models sketch structure, textual appearance, 3D geometry, and multi-view texture evidence. Specifically, the framework distills high-level semantics for sketch encoding, generates intent-aware sketch-text query tokens, extracts geometry and texture representations for each candidate 3D model, and performs query-aware evidence aggregation with fine-grained late-interaction matching. Experiments on ST-Tex3D validate the effectiveness of the proposed task formulation, dataset, and retrieval framework for texture- and style-sensitive 3D model retrieval. Our code and models will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.