SMS enabled fraud is of great concern globally. Building classifiers based on machine learning for SMS fraud requires the use of suitable datasets for model training and validation. Most research has centred on SMS datasets in English. Chichewa is a major language in Africa and is the language of communication in Malawi. This paper introduced a first dataset for SMS fraud detection in Chichewa and reported on experiments with machine learning algorithms for classifying SMSs as fraud or non-fraud. We answered the broader research question of how to develop machine learning classification models for Chichewa SMSs. To do that, we created three datasets. A dataset of SMSs in Chichewa was collected through primary research from a segment of the young population in Blantyre city in Malawi. We applied label-preserving text transformations to increase its size. The enlarged dataset was translated into English using two approaches: human translation and machine translation. The Chichewa and the translated datasets were used to train classification algorithms using random forest, logistic regression, naïve Bayes and support vector machine. The models achieved a promising accuracy of over 96% on the Chichewa dataset. We rerun the machine learning models on the English datasets obtained by translation. The models had a drop in performance when moving from the Chichewa to the machine translated dataset. This highlighted the importance of data preprocessing, especially in multilingual or cross-lingual NLP tasks, and the unsuitability of machine-translation for training models for text classification. Our results underscored the importance of developing language specific models for SMS fraud detection to optimise accuracy and performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Using Machine Learning to Detect Fraudulent SMSs in Chichewa

  • Amelia Taylor,
  • Amoss Robert

摘要

SMS enabled fraud is of great concern globally. Building classifiers based on machine learning for SMS fraud requires the use of suitable datasets for model training and validation. Most research has centred on SMS datasets in English. Chichewa is a major language in Africa and is the language of communication in Malawi. This paper introduced a first dataset for SMS fraud detection in Chichewa and reported on experiments with machine learning algorithms for classifying SMSs as fraud or non-fraud. We answered the broader research question of how to develop machine learning classification models for Chichewa SMSs. To do that, we created three datasets. A dataset of SMSs in Chichewa was collected through primary research from a segment of the young population in Blantyre city in Malawi. We applied label-preserving text transformations to increase its size. The enlarged dataset was translated into English using two approaches: human translation and machine translation. The Chichewa and the translated datasets were used to train classification algorithms using random forest, logistic regression, naïve Bayes and support vector machine. The models achieved a promising accuracy of over 96% on the Chichewa dataset. We rerun the machine learning models on the English datasets obtained by translation. The models had a drop in performance when moving from the Chichewa to the machine translated dataset. This highlighted the importance of data preprocessing, especially in multilingual or cross-lingual NLP tasks, and the unsuitability of machine-translation for training models for text classification. Our results underscored the importance of developing language specific models for SMS fraud detection to optimise accuracy and performance.