<p>Spam detection is crucial for email security. It protects individuals and organizations from unsolicited and potentially malicious emails, which can lead to significant financial losses. Existing spam email detection techniques struggle with the high volume, complexity, and variability of natural language and require extensive feature engineering, resulting in poor performance, including misclassifying legitimate emails as spam, missing some spam emails, and the inability to detect new types of spam emails. This paper proposes a novel spam email detection mechanism based on XLNet, a pre-trained language model that has not been previously explored for this purpose. By fine-tuning XLNet on a labeled dataset of spam and non-spam emails without requiring hand-engineered features, our innovative approach streamlines the detection process. The model predicts the class of previously unseen emails by leveraging the fine-tuned XLNet architecture. The proposed model is evaluated on various benchmark datasets and compared with state-of-the-art models. The model outperforms or is at least comparable to the state-of-the-art models, achieving accuracy, area under the receiver operating characteristic curve (AUC), and F1 scores of 0.9869, 0.9817, and 0.9869 on the SpamAssassin dataset; 0.9892, 0.9893, and 0.9892 on the Enron dataset; 0.9944, 0.9967, and 0.9944 on the Ling-Spam dataset; and 0.9888, 0.9889, and 0.9888 on the combined dataset, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An accurate spam email detection mechanism using XLNet

  • Neeraj Shrestha,
  • Jared Oluoch,
  • Weiqing Sun,
  • Junghwan Kim,
  • Eralda Caushaj

摘要

Spam detection is crucial for email security. It protects individuals and organizations from unsolicited and potentially malicious emails, which can lead to significant financial losses. Existing spam email detection techniques struggle with the high volume, complexity, and variability of natural language and require extensive feature engineering, resulting in poor performance, including misclassifying legitimate emails as spam, missing some spam emails, and the inability to detect new types of spam emails. This paper proposes a novel spam email detection mechanism based on XLNet, a pre-trained language model that has not been previously explored for this purpose. By fine-tuning XLNet on a labeled dataset of spam and non-spam emails without requiring hand-engineered features, our innovative approach streamlines the detection process. The model predicts the class of previously unseen emails by leveraging the fine-tuned XLNet architecture. The proposed model is evaluated on various benchmark datasets and compared with state-of-the-art models. The model outperforms or is at least comparable to the state-of-the-art models, achieving accuracy, area under the receiver operating characteristic curve (AUC), and F1 scores of 0.9869, 0.9817, and 0.9869 on the SpamAssassin dataset; 0.9892, 0.9893, and 0.9892 on the Enron dataset; 0.9944, 0.9967, and 0.9944 on the Ling-Spam dataset; and 0.9888, 0.9889, and 0.9888 on the combined dataset, respectively.