<p>Social media platforms have become susceptible to toxic content, including hate speech and abusive language, posing threats to online communities. Traditional methods like rule-based and regular expression matching have proven inadequate, leading to the adoption of machine learning techniques. While deep learning models have shown promise, their complexity and interpretability remain challenges. This research proposes a novel hybrid strategy for optimizing text classification accuracy in detecting toxic content in Arabic tweets, specifically addressing the scarcity of high-quality labeled datasets. The approach combines Enhanced Non-Dominated Sorting Genetic Algorithm II (NSGA-II) and XGBoost techniques for feature selection, maximising classification accuracy while minimizing the number of features selected. The proposed model utilizes the L-HSAB dataset, a publicly available collection of Arabic Levantine tweets categorized as normal, abusive, or hate speech. The model employs a multi-objective optimization approach, maximizing the representation of selected features and minimizing the diversity of term-category associations. A hybrid algorithm combining NSGA-II and Tabu search is used to streamline the classification process by reducing the number of features. The proposed model demonstrates significant performance improvements over baseline and deep learning solutions, highlighting the importance of feature selection in achieving high prediction performance and computational efficiency. This research contributes a novel approach for detecting toxic content in Arabic tweets, addressing the challenges of limited datasets and interpretability. The hybrid strategy incorporating NSGA-II and XGBoost techniques provides a valuable tool for optimizing text classification accuracy, making it applicable to various data mining problems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing text classification accuracy: a hybrid strategy incorporating enhanced NSGA-II and XGBoost techniques for feature selection

  • Mohamed Atef Mosa

摘要

Social media platforms have become susceptible to toxic content, including hate speech and abusive language, posing threats to online communities. Traditional methods like rule-based and regular expression matching have proven inadequate, leading to the adoption of machine learning techniques. While deep learning models have shown promise, their complexity and interpretability remain challenges. This research proposes a novel hybrid strategy for optimizing text classification accuracy in detecting toxic content in Arabic tweets, specifically addressing the scarcity of high-quality labeled datasets. The approach combines Enhanced Non-Dominated Sorting Genetic Algorithm II (NSGA-II) and XGBoost techniques for feature selection, maximising classification accuracy while minimizing the number of features selected. The proposed model utilizes the L-HSAB dataset, a publicly available collection of Arabic Levantine tweets categorized as normal, abusive, or hate speech. The model employs a multi-objective optimization approach, maximizing the representation of selected features and minimizing the diversity of term-category associations. A hybrid algorithm combining NSGA-II and Tabu search is used to streamline the classification process by reducing the number of features. The proposed model demonstrates significant performance improvements over baseline and deep learning solutions, highlighting the importance of feature selection in achieving high prediction performance and computational efficiency. This research contributes a novel approach for detecting toxic content in Arabic tweets, addressing the challenges of limited datasets and interpretability. The hybrid strategy incorporating NSGA-II and XGBoost techniques provides a valuable tool for optimizing text classification accuracy, making it applicable to various data mining problems.