Phishing Email Clustering Using Quasi-Ground-Truth Machine Learning
Abstract
Fast and precise detection of new and unknown phishing campaigns is essential for incident response teams and, more broadly, for overall cybersecurity. However, no reliable automated phishing detection methods exist, largely due to varying incident clustering criteria and the scarcity of ground-truth datasets. We address the problem by providing methods for creating quasi-ground truth, enabling the development of two machine-learning clustering models with a Siamese architecture. The models, trained on a new dataset of phishing email, prove transferable onto a test dataset of a different origin. While the generalization of the model's application is not fully satisfactory, we find that it may be improved through cooperation between the two models. The new phishing dataset has been released publicly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.