Promptable Geometric Grounding in B-Rep Models
Abstract
We study the task of grounding in Boundary Representation (B-Rep) models, a foundational capability for AI-driven Computer-aided design (CAD) modeling. Such grounding provides a semantic interface for expressing user intent, yet remains largely unexplored. Unlike traditional 3D visual grounding that localizes objects in real-world point cloud scenes, the task requires both a fundamentally different scene representation and a finer localization granularity, where targets are not objects but topologically defined primitives (e.g., faces and edges) embedded in a structured geometric model. This paper introduces PGGM, a Promptable Geometric Grounding Model for B-Rep models. PGGM jointly supports natural language descriptions and partial selections as prompts. It decouples continuous geometry and discrete topology: pretrained VAEs encode geometric shape information, while a topology encoder captures structured primitive relations. These representations are then integrated through prompt-conditioned fusion and primitive-level decoding for fine-grained grounding. To further improve hard-negative discrimination, we introduce Discriminative Bootstrapping (DB), a training strategy that leverages the partial branch to turn noisy auxiliary signals into discriminative supervision. DB constructs complementary signals from external teachers, model self-bootstrapping, and proximity-based hard negatives, exposing the model to distractors that are semantically plausible, self-predicted, or close to the target in latent space. This enables more reliable grounding under fine-grained geometric ambiguity. Experiments show that PGGM significantly improves B-Rep grounding over strong baselines, achieving a +17.49% F1 gain and a +24.18% precision gain in the text-only setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.