Generative-Discriminative Knowledge Distillation for Cross-Lingual NER
Abstract
Cross-lingual named entity recognition (NER) aims to build an NER model that generalizes to the low-resource target languages with labeled data from the high-resource source language. Existing approaches can be broadly categorized into two paradigms: discriminative methods based on small language models and generative methods based on large language models. These paradigms are complementary: discriminative models enable fast and efficient inference but may overfit source-language surface features, whereas generative models are generally more robust to unseen languages but incur substantial computational costs during decoding. To combine their complementary strengths, we propose a **Gen**erative-**Dis**criminative **K**nowledge **D**istillation **(GenDisKD)** framework, which transfers knowledge from both discriminative and generative teachers to a compact student model for cross-lingual NER. Furthermore, we introduce a candidate-set distillation loss that preserves alternative candidate labels when teachers disagree, allowing the student to resolve ambiguities through complementary supervision. Extensive experiments on three benchmarks demonstrate that GenDisKD consistently outperforms existing discriminative and generative methods, validating its effectiveness and robustness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.