BERT-Based Model for Sarcasm Detection in Arabic Texts
摘要
In the current digital era, social networks have become privileged spaces of expression for individuals, allowing them to share their opinions. However, the informal and often ambiguous nature of communications on these platforms can make it difficult to interpret the tone and intent of messages. In particular, the form of irony, sarcasm, can be particularly tricky to pin down. The automatic detection of sarcasm in Arabic texts is a recent area of research that has experienced significant growth in recent years. Indeed, understanding sarcasm is essential for an accurate analysis of feelings and opinions expressed on social networks. Our article is part of this dynamic by presenting our research work on the detection of sarcasm in tweets in dialectal Arabic. We developed a corpus of annotated data, composed of 19669 tweets labeled according to their sarcastic polarity. This corpus was then used to train and evaluate different classification models based on the BERT (Bidirectional Encoder Representations from Transformers) architecture. Our results indicate that the Arabert-V2-twitter model, finely tuned to our corpus of data, achieves an accuracy score of 81.24% for sarcasm classification.