Multi-ideology Dataset for Detecting Online Extremism in Kazakh Language
摘要
The rapid growth of Internet and social media users has led to a parallel increase in the use of these platforms by extremists to spread ideologies, radicalize individuals, and recruit members. This poses significant challenges for maintaining the security of online spaces and necessitates the development of effective detection strategies. This study addresses the problem of identifying extremist content in the Kazakh language by creating a high-quality dataset specifically tailored to this need. The dataset was compiled from multiple sources, including social networks, news articles, and academic publications, and classified into four categories: radicalization, propaganda, recruitment, and neutral texts. This dataset is an essential tool for researchers and practitioners aiming to develop machine learning models that can automatically detect and classify extremist messages, ultimately contributing to the prevention of online radicalization and extremism.