<p>In today’s advanced digital technology, fraudulent messages frequently deceive individuals. Many users unknowingly engage with links shared by cybercriminals, unaware of the potential risks. Such impulsive actions often lead to financial losses or other adverse consequences. Beyond educating users, it is crucial to identify and filter out spam messages effectively. While various machine and deep learning-based approaches exist, they are primarily limited to one or two languages, particularly English, and rely heavily on contextual feature extraction. This paper presents an innovative framework for spam message classification using non-contextual meta-data features across multiple Indian languages. Experiments were conducted on six diverse, multilingual SMS datasets. The extracted features were processed using a Random Forest (RF) classifier and validated through Recursive Feature Elimination (RFE) and eXplainable Artificial Intelligence (XAI) techniques. The proposed framework achieved 94–97% accuracy, demonstrating its effectiveness compared to state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

NCMDFE: performance evaluation of non-contextual meta-data feature extraction for spam SMS classification in Indian Languages

  • Ramanujam E,
  • Abirami AM

摘要

In today’s advanced digital technology, fraudulent messages frequently deceive individuals. Many users unknowingly engage with links shared by cybercriminals, unaware of the potential risks. Such impulsive actions often lead to financial losses or other adverse consequences. Beyond educating users, it is crucial to identify and filter out spam messages effectively. While various machine and deep learning-based approaches exist, they are primarily limited to one or two languages, particularly English, and rely heavily on contextual feature extraction. This paper presents an innovative framework for spam message classification using non-contextual meta-data features across multiple Indian languages. Experiments were conducted on six diverse, multilingual SMS datasets. The extracted features were processed using a Random Forest (RF) classifier and validated through Recursive Feature Elimination (RFE) and eXplainable Artificial Intelligence (XAI) techniques. The proposed framework achieved 94–97% accuracy, demonstrating its effectiveness compared to state-of-the-art methods.