Emoji Retrieval from Gibberish or Garbled Social Media Text: A Novel Methodology and a Case Study
摘要
Emojis, considered an integral aspect of social media conversations, are widely used on almost all social media platforms. However, social media data may be noisy and may also include gibberish or garbled text which is difficult to detect and work with. Most naïve data preprocessing approaches recommend removing such gibberish or garbled text from social media posts before performing any form of data analysis or before passing such data to any machine learning model. However, it is important to note that such gibberish or garbled text may have been an emoji(s) in the original social media post(s) and failure to retrieve the actual emoji(s) may result in the loss or lack of contextual meaning of the analyzed social media data. The work presented in this paper aims to address this challenge by proposing a three-step reverse engineering-based novel methodology for retrieving emojis from garbled or gibberish text in social media posts. The development of this methodology also helped to unravel the reasons that could lead to the generation of gibberish or garbled text related to data mining of social media posts. To evaluate the effectiveness of the proposed methodology, the model was applied to a dataset of 509,248 Tweets about the Mpox outbreak, that has been used in about 30 prior works in this field, none of which were able to retrieve the emojis in the original Tweets from the gibberish text present in this dataset. Using our methodology, we were able to retrieve a total of 157,748 emojis present in 76,914 Tweets in this dataset by processing the gibberish or garbled text. The effectiveness of this methodology has been discussed in the paper through the presentation of multiple metrics related to text readability and text coherence which include the Flesch Reading Ease, Flesch Kincaid Grade Score, Coleman Liau index, Automated Readability Index, Dale Chall Readability Score, Text Standard, and Reading Time for the Tweets before and after the application of the methodology to the Tweets. The results showed that the application of this methodology to the Tweets improved the readability and coherence scores. Finally, as a case study, the frequency of emoji usage in these Tweets about the Mpox outbreak was analyzed and the results are presented.