Empowering Hate Speech Detection: A Comparative Exploration of Deep Learning Models
摘要
Hate Speech Detection, employing Natural Language Processing (NLP) and Machine Learning techniques, serves as a crucial tool for automatically identifying and flagging discriminatory language, hatred, or other protected characteristics within the vast landscape of social media. The escalating volume of social media posts, generated every second, has led to a surge in hateful and disrespectful comments. The imperative to automatically detect such content arises due to the time and resource intensity involved in creating high-quality human-labeled datasets. This study meticulously selects a dataset ensuring its inclusion of a substantial amount of hate speech instances across diverse languages, capturing the nuanced nature of online discourse. To enhance model performance, each model, including BERT, GPT, and LSTM, undergoes fine-tuning on this dataset, incorporating personalized adjustments based on hyperparameters and preprocessing strategies. The comparative analysis yields insightful findings on the effectiveness of BERT, GPT, and LSTM models in Hate Speech Detection. Notably, this research advances beyond the previous abstract, acknowledging the diverse linguistic landscape by utilizing data generated from social media in various languages. The study contributes valuable insights to intellectuals and researchers, emphasizing the deployment of state-of-the-art techniques in real-world applications across linguistic boundaries. This paper underscores the importance of a nuanced approach in model selection for different Hate Speech Detection tasks, recognizing the challenges posed by diverse languages. Moreover, it highlights the necessity of research-centric innovation in addressing the critical issue of hate speech across multilingual social media platforms.