SEAR: Seeing What Matters via Structure-Guided Language
Abstract
Fine-grained and specialized image classification often relies on subtle differences in visual structure and appearance, including shape, color, texture, and localized defect patterns. Image-level labels provide limited guidance on which cues matter and where they occur. Pretrained vision-language models can make visual attributes explicit in text, but generic captions may not highlight the structural properties that distinguish classes. We introduce SEAR (Semantic Enhancement via Alignment and Reconstruction Network) to make discriminative visual structure explicit in language and fuse the resulting descriptions with image features, helping identify subtle but decisive cues. A frozen VLM generates these descriptions under explicit instructions to focus on parts, shapes, and local patterns. Alignment encourages semantic consistency between the two views, while reconstruction helps preserve complementary visual evidence that the descriptions may omit. A finite-sample analysis characterizes a text-quality threshold: generated text helps when the recoded structural signal outweighs generator noise and the additional estimation cost. Across eight benchmarks, SEAR improves mean F1-macro over the ResNet-101 image-only baseline by 0.32-7.68 percentage points, including 2.43 points on Steel Defect, where text alone performs poorly. On a 20-class CUB subset, SEAR with task-prompted descriptions also outperforms its counterpart using human-authored captions, suggesting that faithfully describing image content may not emphasize the attributes most useful for classification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.