Sentiment Analysis for Moroccan Dialect Using the Model of Machine Learning
摘要
Since the advent of social networks, the scientific community working on NLP has been increasingly interested in developing automatic sentiment analysis and opinion-mining tools. Currently, most of the proposed works deal with Indo-European languages, especially English. However, a large community of people who use dialectics is not yet targeted. In this work, we are interested in the Moroccan Arabic dialect (MAD). Our primary focus, in this sense, is to determine the polarity of a given text statement. For the dataset, we have used data that are collected from different social networks (Facebook, Twitter, YouTube, Instagram, and Websites). Indeed, to our knowledge, there is no publicly available MAD dataset based on all social networks for the sentiment analysis task. Moreover, our collected data dataset is the largest Moroccan dataset created for sentiment analysis. It is characterized by its size, its quality, and its variety. We describe the methodology we evolved for collecting, preprocessing, feature extraction, and polarity classification using various supervised machine learning algorithms.