Prompting with Connections: Graph-Guided Attribute Selection for Test-Time Calibration
Abstract
Test-time prompt tuning adapts vision-language models (VLMs) to unseen target domains without labels by minimizing prediction entropy over augmented views. While boosting zero-shot accuracy, this process often induces overconfidence and degrades calibration. Existing approaches initialize prompts with LLM-generated visual attributes under contrastive regularization, but treat attributes independently, ignoring their relational structure. We introduce ARGTCA, a graph-guided framework for calibrated test-time prompt tuning that explicitly models relationships among class-specific visual attributes. By representing class–attribute pairs as nodes and linking within-class siblings and identical attributes across classes in a Symbolic Attribute Graph, a GAT trained with a supervised contrastive objective produces relational embeddings and makes attribute genericity—how many classes share an attribute—legible as distance from the class centroid (Spearman on Caltech101). The embeddings drive two prompt-initialization strategies: ARGTCA-Div selects diverse, complementary attributes, while ARGTCA-Disc selects highly discriminative ones; the selected attributes then initialize the test-time prompts. Across ten visual-classification benchmarks on CLIP ViT-B/16, ARGTCA-Div reduces average ECE to , a reduction of up to relative to existing test-time prompt-tuning methods, while ARGTCA-DISC attains the highest average accuracy () at a competitive ECE of . On ResNet-50, ARGTCA-DIV again attains the lowest average ECE () and ARGTCA-DISC the second lowest (), confirming generalization across CNN and Transformer encoders. Attribute selection runs once, offline, adding no test-time overhead. The codebase will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.