This paper focuses on the identification and classification of Arabic Verbal Multi-Word Expressions (VMWE), which are combinations of words with a unitary meaning and containing at least one verb. We describe our contribution to annotating the Arabic-annotated-conll17 corpus with VMWEs, following the annotation guide of the PARSEME framework, and using the ChatGPT model to enrich the corpus with sentences that contain more linguistic phenomena. We propose a method for the identification and classification of Arabic VMWEs based on machine learning (Random Forest, Decision trees and Support Vector Machine) and deep learning (Conv1D+BiLSTM) models. The results show that our Random Forest model outperformed others in identification, while the Decision Tree model excelled in classification. Additionally, our deep learning model exhibited significant performance improvement after data enrichment with ChatGPT. These findings underscore the effectiveness of our approach in addressing this complex problem.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Learning Approach to Identify and Classify Arabic Verbal Multi-word Expressions

  • Meniara Bsissa,
  • Iskandar Keskes,
  • Najet Hadj Mohamed

摘要

This paper focuses on the identification and classification of Arabic Verbal Multi-Word Expressions (VMWE), which are combinations of words with a unitary meaning and containing at least one verb. We describe our contribution to annotating the Arabic-annotated-conll17 corpus with VMWEs, following the annotation guide of the PARSEME framework, and using the ChatGPT model to enrich the corpus with sentences that contain more linguistic phenomena. We propose a method for the identification and classification of Arabic VMWEs based on machine learning (Random Forest, Decision trees and Support Vector Machine) and deep learning (Conv1D+BiLSTM) models. The results show that our Random Forest model outperformed others in identification, while the Decision Tree model excelled in classification. Additionally, our deep learning model exhibited significant performance improvement after data enrichment with ChatGPT. These findings underscore the effectiveness of our approach in addressing this complex problem.