Aspect-Category-Opinion-Sentiment Quadruple Extraction for Code Mixed Data
摘要
Over the years, researchers have delved deeply into Aspect-Based Sentiment Analysis (ABSA), advancing techniques like Aspect Sentiment Triplet Extraction (ASTE) and Aspect-Category-Opinion-Sentiment Quadruple Extraction (ACOSQE). Since traditional sentiment analysis techniques frequently fail to capture the nuanced opinions of different aspects within a sentence, the need for Aspect-Based Sentiment Analysis grows. While much of this exploration has focused on languages like English and Chinese, as well as code-mixed datasets like Chinese-English, there remains a notable gap in the research when it comes to Hindi-English datasets with ACOS quadruples. This project aims to fill that void by creating a novel dataset in Hinglish, based on mobile phone reviews. Comprehensive annotation guidelines and a framework were designed to carefully annotate the dataset. The newly curated data was validated using state-of-the-art statistical analysis methodologies to ensure inter-annotator agreement on subjective aspects of the dataset, i.e., sentiment. Furthermore, the model is evaluated using the Extract-Classify-ACOS model on a multilingual vocabulary to benchmark the dataset.