Religious Sentiment Analysis and Detection on Social Media
摘要
This research addresses the rising concern of religious hate speech on Facebook, employing a comprehensive methodology to create a dataset by annotating statements from diverse religious backgrounds. After preprocessing steps like noise removal and tokenization, various linguistic features, including character N-grams, TF-IDF-weighted N-grams, and text quality indicators, were utilized. Multiple supervised learning algorithms, such as logistic regression, linear SVC, random forest, Naive Bayes, decision tree, and a GRU-based neural network model, were examined for optimal performance. Evaluation metrics like accuracy, precision, recall, and F1 score were employed. The study illuminates the prevalence and characteristics of religiously motivated hate speech on social media, emphasizing the importance of algorithm selection based on desired performance metrics. This research offers a valuable tool for academics and platform management to effectively identify and mitigate the impact of religious hate speech on digital platforms, supporting ongoing efforts to combat online discrimination.