Iraqi Dialect Emoji Prediction Based on Deep CNN-LSTM Architecture
摘要
Emojis have become an essential part of online communication as social media platforms have grown in popularity. These small graphical symbols express a broad range of emotions and sentiments and represent objects or concepts. The widespread use of emojis has piqued the interest of researchers because they convey valuable semantic and sentiment information to supplement textual content visually. Predicting Iraqi dialect emojis is difficult due to a lack of labelled data and the complexities associated with understanding the cultural nuances associated with these emojis. The study employed NodeXL over three months to collect over 4 million tweets from X (formerly Twitter), initialising the raw dataset through preprocessing. An architecture combining the convolutional neural network (CNN) and long short-term memory (LSTM) network can predict emojis in Arabic tweets and comments using the Iraqi dialect, a slang-based Arabic variety, to effectively predict the occurrence of emojis in the text. The proposed models in this paper (LSTM, bidirectional LSTM (BiLSTM), BiLSTM attention, and CNN-LSTM) achieved impressive accuracy rates of 94.01%, 94.1%, 93.39%, and 96.15%, respectively, in emoji classification tailored explicitly for the Iraqi dialect, outperforming previous research efforts in this domain. This study proposes a method for predicting emojis in Arabic tweets and comments, particularly in the slang-based Arabic variety.