Spam and ham detection is a critical function of email and SMS filtering systems that prevents unwanted and even dangerous messages from reaching users’ inboxes. This research focuses on developing exact ways for discriminating between ham (legitimate, non-spam messages) and spam (unsolicited bulk messages) in order to enhance email and SMS security and user experience. The study analyses various machine learning algorithms, using elements including message content (such as SMS and email), metadata, and sender information in order to achieve excellent performance in the categorization task. The dimensionality reduction approach allowed for the retrieval of significant features. The data was subjected to a number of cutting-edge machine learning classifiers, including decision trees, support vector machines (SVMs), random forests (RFs), and ensembling algorithms. Performance is evaluated using metrics such as accuracy, precision, recall, and F1-Score, with a focus on finding a balance between raising detection rates and lowering false positives.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Ensemble Methods for Spam Classification

  • Y. Shashikala,
  • R. Satyasri,
  • S. Shaheena,
  • Padma Jyothi Uppalapati,
  • Adina Karunasri,
  • K. Ratna Kumari

摘要

Spam and ham detection is a critical function of email and SMS filtering systems that prevents unwanted and even dangerous messages from reaching users’ inboxes. This research focuses on developing exact ways for discriminating between ham (legitimate, non-spam messages) and spam (unsolicited bulk messages) in order to enhance email and SMS security and user experience. The study analyses various machine learning algorithms, using elements including message content (such as SMS and email), metadata, and sender information in order to achieve excellent performance in the categorization task. The dimensionality reduction approach allowed for the retrieval of significant features. The data was subjected to a number of cutting-edge machine learning classifiers, including decision trees, support vector machines (SVMs), random forests (RFs), and ensembling algorithms. Performance is evaluated using metrics such as accuracy, precision, recall, and F1-Score, with a focus on finding a balance between raising detection rates and lowering false positives.