SecEmbedEval: A Comprehensive Benchmark for Evaluating Text Embeddings in Cybersecurity
Abstract
We introduce SecEmbedEval, a benchmark for evaluating cybersecurity text embeddings. SecEmbedEval contains over 17,000 task instances in Chinese and English, built from public cybersecurity sources through source-aware preprocessing, large language model (LLM)-assisted construction, and blind verification. It covers nine tasks in five categories: semantic similarity, retrieval, classification, clustering, and efficiency. To support aggregated evaluation, we propose two metrics. SEE-Score computes a min-max normalized average across five task measures. Transfer Discordance (TD) quantifies the share of model pairs whose ranking order differs between public leaderboards and our in-domain benchmark. We evaluate 12 representative embedding models. General-domain standing does not transfer (TD ). Frozen cybersecurity-pretrained encoders trail contrastively trained general-purpose models even under English-only controls. No evaluated model, including the strongest commercial system, can ground bare CVE identifiers to their advisories, and clustering remains challenging for every model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.