<p>Sentiment analysis plays a crucial role in understanding public opinion across Arabic social media platforms; however, challenges such as morphological richness, dialectal diversity, and contextual ambiguity hinder model performance. This study evaluates the impact of annotation quality and model architecture on the accuracy of Arabic sentiment analysis by comparing traditional machine learning algorithms with transformer-based architectures. A novel dataset was developed using an active annotation strategy that combines human expertise and artificial intelligence (AI) assisted labelling through Chat Generative Pre-Trained Transformer versions 3.5 and 4. This study investigates how annotation quality affects classification performance metrics —accuracy, precision, recall, and F1-score— across various classifiers. The experimental results reveal that annotation consistency exerts a stronger influence on performance than model complexity. Classical machine learning models such as k nearest neighbours, multinomial naïve Bayes and random forest performed best with ChatGPT 4-labelled data, reflecting the lexical coherence of AI generated annotations. Conversely, context-sensitive models including support vector machine and logistic regression, achieved superior results with human-labelled datasets, emphasizing the value of human semantic interpretation. Transformer-based architectures—AraBERTv2, CAMeL Lab BERT, and XLM-RoBERTa—outperformed all other models, demonstrating robust contextual understanding and resilience to annotation noise. The findings indicate that integrating AI-assisted annotation with transformer-based deep learning provides a scalable and linguistically adaptive framework for Arabic sentiment analysis. This study advances both theoretical understanding and practical applications by proposing a balanced human–artificial intelligence annotation paradigm that enhances the efficiency and interpretability of Arabic sentiment analysis models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Arabic sentiment analysis during the Qatar world cup: a study of generative pre-trained transformers and machine learning techniques

  • Ghaleb Al-Gaphari,
  • Salah AL-Hagree,
  • Maher Al-Sanabani,
  • Ahmed Al-Shalabi,
  • Baligh Al-Helali,
  • Ala’a Alawdi

摘要

Sentiment analysis plays a crucial role in understanding public opinion across Arabic social media platforms; however, challenges such as morphological richness, dialectal diversity, and contextual ambiguity hinder model performance. This study evaluates the impact of annotation quality and model architecture on the accuracy of Arabic sentiment analysis by comparing traditional machine learning algorithms with transformer-based architectures. A novel dataset was developed using an active annotation strategy that combines human expertise and artificial intelligence (AI) assisted labelling through Chat Generative Pre-Trained Transformer versions 3.5 and 4. This study investigates how annotation quality affects classification performance metrics —accuracy, precision, recall, and F1-score— across various classifiers. The experimental results reveal that annotation consistency exerts a stronger influence on performance than model complexity. Classical machine learning models such as k nearest neighbours, multinomial naïve Bayes and random forest performed best with ChatGPT 4-labelled data, reflecting the lexical coherence of AI generated annotations. Conversely, context-sensitive models including support vector machine and logistic regression, achieved superior results with human-labelled datasets, emphasizing the value of human semantic interpretation. Transformer-based architectures—AraBERTv2, CAMeL Lab BERT, and XLM-RoBERTa—outperformed all other models, demonstrating robust contextual understanding and resilience to annotation noise. The findings indicate that integrating AI-assisted annotation with transformer-based deep learning provides a scalable and linguistically adaptive framework for Arabic sentiment analysis. This study advances both theoretical understanding and practical applications by proposing a balanced human–artificial intelligence annotation paradigm that enhances the efficiency and interpretability of Arabic sentiment analysis models.