In the current digital era, social networks have become privileged spaces of expression for individuals, allowing them to share their opinions. However, the informal and often ambiguous nature of communications on these platforms can make it difficult to interpret the tone and intent of messages. In particular, the form of irony, sarcasm, can be particularly tricky to pin down. The automatic detection of sarcasm in Arabic texts is a recent area of research that has experienced significant growth in recent years. Indeed, understanding sarcasm is essential for an accurate analysis of feelings and opinions expressed on social networks. Our article is part of this dynamic by presenting our research work on the detection of sarcasm in tweets in dialectal Arabic. We developed a corpus of annotated data, composed of 19669 tweets labeled according to their sarcastic polarity. This corpus was then used to train and evaluate different classification models based on the BERT (Bidirectional Encoder Representations from Transformers) architecture. Our results indicate that the Arabert-V2-twitter model, finely tuned to our corpus of data, achieves an accuracy score of 81.24% for sarcasm classification.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BERT-Based Model for Sarcasm Detection in Arabic Texts

  • Amal Mezghani,
  • Rahma Boujelben,
  • Mariem Ellouze

摘要

In the current digital era, social networks have become privileged spaces of expression for individuals, allowing them to share their opinions. However, the informal and often ambiguous nature of communications on these platforms can make it difficult to interpret the tone and intent of messages. In particular, the form of irony, sarcasm, can be particularly tricky to pin down. The automatic detection of sarcasm in Arabic texts is a recent area of research that has experienced significant growth in recent years. Indeed, understanding sarcasm is essential for an accurate analysis of feelings and opinions expressed on social networks. Our article is part of this dynamic by presenting our research work on the detection of sarcasm in tweets in dialectal Arabic. We developed a corpus of annotated data, composed of 19669 tweets labeled according to their sarcastic polarity. This corpus was then used to train and evaluate different classification models based on the BERT (Bidirectional Encoder Representations from Transformers) architecture. Our results indicate that the Arabert-V2-twitter model, finely tuned to our corpus of data, achieves an accuracy score of 81.24% for sarcasm classification.