Behavior-Based Insider Threat Detection Using Machine Learning: A Comparative Study of Isolation Forest, Random Forest, and XGBoost
Abstract
Detecting insider threats from enterprise activity data is challenging because malicious behavior can occur under legitimate credentials and often overlaps with normal user activity. We study how behavioral representation and learning strategy affect insider-threat detection under severe class imbalance. Using the CERT Insider Threat Dataset r4.2, we transform heterogeneous logon, device, file, and email activity into 330,452 user-day observations and construct 49 behavioral features capturing activity volume, temporal patterns, email behavior, computer usage, and deviations from user-specific historical behavior. We evaluate three complementary learning paradigms: Isolation Forest as an unsupervised anomaly detector, Random Forest as a supervised ensemble classifier, and XGBoost as a supervised gradient-boosting model. To reduce temporal information leakage and better reflect deployment conditions, we use chronological training, validation, and held-out test partitions, with model selection and threshold optimization performed exclusively on the validation period. On the held-out test set, XGBoost using all 49 behavioral features achieves 79.23% precision, 59.43% recall, 67.92% F1-score, 99.57% ROC-AUC, and 78.97% PR-AUC. Compared with an 11-feature representation, the 49-feature representation substantially improves XGBoost performance, increasing F1-score from 34.40% to 67.92% and PR-AUC from 24.48% to 78.97%. The comparison further reveals a trade-off between unsupervised anomaly detection and supervised classification: Isolation Forest attains 57.38% recall but only 0.97% precision, while Random Forest achieves 99.48% accuracy but only 15.57% recall. These results indicate that richer behavioral representations incorporating temporal and user-specific historical context can be more consequential than model selection alone for detecting rare insider-threat behavior. We additionally use SHAP to provide feature-level explanations for XGBoost predictions, linking detected threats to interpretable behavioral deviations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.