Improving Arabic NER in Social Media: Performance Analysis of Pre-trained Models and Data Augmentation
摘要
Named Entity Recognition (NER) plays a vital role in extracting structured information from unstructured text, especially on social media platforms such as Twitter. Deep neural networks are effective for a variety of natural language processing tasks; nevertheless, they often require large sets of annotated data sets to outperform simpler models. This data may not be sufficiently diverse or available, and collecting and annotating them can be a time-consuming and costly process. The aim of this paper is to study the impact of data augmentation on the Named Entity Recognition (NER) task on a text issue from social media, using pre-trained models in Arabic. Two scenarios were set up to evaluate the models, starting with a scenario without data augmentation, in which the MARBERT v2 model surpassed the other models with an F1 score of 67.4%. Then, the second scenario with data augmentation, in which the Arabert v0.2 twitter model obtained the best F1 score with 72.16%.