Benchmarking sentiment analysis of algerian arabic dialect on X (Twitter)
摘要
The rise of Arabic dialects on social media has made sentiment analysis essential for understanding regional opinions. However, most existing research focuses on high-resource languages, leaving low-resource dialects like Algerian Arabic underexplored. To address this, we introduced a publicly available dataset of 18,589 Algerian dialect tweets and presented a customized preprocessing pipeline tailored to the dialect’s unique linguistic features. We further conducted a comprehensive evaluation of sentiment analysis models, where our best-performing MARBERT-LSTM model achieved 91.23% accuracy, setting a new benchmark. This work provides both a valuable resource and a strong baseline for future research in dialectal Arabic NLP.