This paper evaluates sentiment quantification methods applied to Brazilian Portuguese corpora. Sentiment quantification, distinct from sentiment classification, estimates the distribution of sentiment classes (positive and negative) within a dataset. We investigate several quantification techniques, including the family Classify and Count (CC) and more sophisticated methods, such as Kernel Density Estimation (KDE) and Distribution y-Similarity (DyS). Our analysis uses five datasets, each containing different distributions of sentiment classes. Our experimental results indicate that KDE and DyS methods consistently outperform others, achieving the best average ranks in terms of quantification accuracy. Statistical tests, including the Friedman and Nemenyi tests, confirm significant performance differences among the methods, with KDE and DyS showing statistically significant improvements over the baseline CC method. These findings highlight the importance of choosing robust quantification techniques for accurate sentiment quantification in corpora across different domains.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Sentiment Quantification Methods in Brazilian Portuguese Corpora

  • Lucas Nildaimon dos Santos Silva,
  • Diego Furtado Silva,
  • Helena de Medeiros Caseli

摘要

This paper evaluates sentiment quantification methods applied to Brazilian Portuguese corpora. Sentiment quantification, distinct from sentiment classification, estimates the distribution of sentiment classes (positive and negative) within a dataset. We investigate several quantification techniques, including the family Classify and Count (CC) and more sophisticated methods, such as Kernel Density Estimation (KDE) and Distribution y-Similarity (DyS). Our analysis uses five datasets, each containing different distributions of sentiment classes. Our experimental results indicate that KDE and DyS methods consistently outperform others, achieving the best average ranks in terms of quantification accuracy. Statistical tests, including the Friedman and Nemenyi tests, confirm significant performance differences among the methods, with KDE and DyS showing statistically significant improvements over the baseline CC method. These findings highlight the importance of choosing robust quantification techniques for accurate sentiment quantification in corpora across different domains.