acceptodds
Under review as a conference paper at ICLR 2027

StructAlign: Learning What to Preserve and What to Complement for Sketch-Text 3D Retrieval

Abstract

Retrieving textured 3D objects from freehand sketches and text requires combining complementary yet unequally reliable cues. Sketches provide strong geometric constraints through contours, proportions, and part relationships, whereas text supplies category and appearance information but may be incomplete, ambiguous, or inconsistent with the sketch. Existing fusion methods often lack explicit protection of sketch structure, allowing textual semantics to distort geometry-sensitive representations. We present SketchAnchor, which treats the sketch representation as a geometric anchor and uses multimodal large language model features to complement missing semantics. It constructs a foreground-aware sketch representation and selectively injects locally relevant semantic features through cross-modal attention and gated residual updates. A foreground Structural KL constraint limits changes in the relative importance of key stroke regions. To align semantic completion with downstream retrieval, we further introduce a shape–texture graded ranking reward and retrieval-drift-calibrated GRPO with completion-specific adaptive clipping. We also construct ST-Tex3D v2, containing 15,600 textured 3D instances with aligned sketches, retrieval-oriented descriptions and texture variants. Experiments demonstrate the effectiveness of explicitly modeling asymmetric modality reliability for fine-grained multimodal retrieval

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.