acceptodds
Under review as a conference paper at ICLR 2027

DualDistill: Transferring Generative Recognition into Efficient Discriminative Prediction for Annotation-Free Multi-Label Visual Recognition

Abstract

Traditional multi-label visual recognition relies heavily on human-annotated data, making it costly to scale to large datasets and label vocabularies. Multimodal Large Language Models (MLLMs) provide a promising alternative by performing annotation-free recognition through autoregressive label generation. However, this generative paradigm incurs substantial inference overhead due to task-prompt encoding and token-by-token decoding. In this paper, we investigate how to transfer the generative recognition capability of MLLMs into flexible and efficient discriminative prediction. We identify two key challenges in this transition: removing the task-aware prompt discards valuable candidate-label semantics, while directly optimizing MLLMs with a discriminative classification objective introduces a mismatch with their original generative training paradigm. To address these challenges, we propose DualDistill, a progressive generation-to-discrimination framework consisting of Cross-Paradigm Distillation (CPD) and Cross-Token Distillation (CTD). CPD first converts generative recognition into prompt-aware discriminative prediction while preserving the generative learning signal. CTD then transfers the resulting prompt-aware capability to an image-only pathway for efficient inference. Extensive experiments on multiple benchmarks demonstrate that DualDistill substantially improves both recognition performance and inference efficiency over strong MLLM fine-tuning and annotation-free baselines, achieving up to a 14 times inference speedup while maintaining strong classification performance without requiring any human annotations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.