Skill-RAG: Retrieval-augmented LLMs for Extreme Multilabel Skill Classification of Unstructured Job Ad Texts
Abstract
Skill Classification, the process of identifying skill-related information from job ads and mapping it to a fine-grained set of skill labels, opens up new possibilities for social science research on labor market dynamics. Mainstream skill taxonomies are generally large-scale, such as the ESCO framework with 13,939 labels, making this an extreme multi-label classification (XMLC) task. While Large Language Models (LLMs) prompted with job ads and skill ontologies offer robust automated classification capabilities, deploying them on massive, fine-grained taxonomies introduces excessive inputs, thereby triggering issues of context window constraints, long-context performance degradation, and positional bias. To overcome these challenges, we propose Skill-RAG, a novel pipeline that integrates a Retrieval-Augmented Generation (RAG) framework into the inference process for XMLC. Specifically, we systematically explore various knowledge base construction paradigms, query granularity strategies (raw vs. span-based retrieval), and multi-stage retrieval pipelines to optimize retrieval accuracy. Furthermore, for downstream LLM-based classification, we comprehensively evaluate prompting and reasoning strategies across three core dimensions: label space cardinality, label ordering, and thought-guided reasoning. Our empirical results demonstrate that all aforementioned components are critical to final classification quality, and our proposed pipeline substantially outperforms previous state-of-the-art methods on standard benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.