<p>Sentiment analysis is the process of determining the expressive direction of the user reviews. Recently, sentiment analysis gets more attention. However, low data sentiment analysis receives less attention. The existing works try to augment the samples to consider this issue. In this study, we have utilized a semi-supervised approach to propose a new approach for low-data sentiment analysis. To do so, we have utilized pre-trained XLNet as a feature extractor network to initialize the feature vector for each tweet. Next, these initial representations are fed into the embedding update module to map features into the new space by optimizing the contrastive loss. Then, we utilized a semi-supervised boosting method to assign pseudo labels to unlabeled data. The iteration between the semi-supervised module and the embedding update module is done until convergence is happened. During these iterations, the embedding update module propagates the error-correcting signals to a semi-supervised module. To evaluate the proposed approach, we have applied it to the SemEval2017dataset (task 4), Sentiment 140, and IMDB Movie Reviews. We have designed many different experiment settings to validate the proposed approach’s different modules. On SemEval2017dataset (task 4), we have got 75.9% and 77.1% in AvgRec and <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10489_2024_6071_Article_IEq1.gif" Format="GIF" Height="21" Rendition="HTML" Resolution="72" Type="Linedraw" Width="36" /> </InlineMediaObject> <EquationSource Format="TEX">\({F}_{1}^{PN}\)</EquationSource> <EquationSource Format="MATHML"><math> <msubsup> <mi>F</mi> <mrow> <mn>1</mn> </mrow> <mrow> <mi mathvariant="italic">PN</mi> </mrow> </msubsup> </math></EquationSource> </InlineEquation> respectively. Also, when only 10% of the training samples as labeled samples are used, we get the 71.8% and 73.6% in AvgRec and <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10489_2024_6071_Article_IEq2.gif" Format="GIF" Height="21" Rendition="HTML" Resolution="72" Type="Linedraw" Width="36" /> </InlineMediaObject> <EquationSource Format="TEX">\({F}_{1}^{PN}\)</EquationSource> <EquationSource Format="MATHML"><math> <msubsup> <mi>F</mi> <mrow> <mn>1</mn> </mrow> <mrow> <mi mathvariant="italic">PN</mi> </mrow> </msubsup> </math></EquationSource> </InlineEquation> respectively. The results show that our approach significantly improves with respect to the comparable methods. Also, on IMDB Movie Reviews and Sentiment 140, the proposed approach demonstrates improved performance compared to comparable methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SSSA: low data sentiment analysis using boosting semi-supervised approach and deep feature learning network

  • Shima Rashidi,
  • Jafar Tanha,
  • Arash Sharifi,
  • Mehdi Hosseinzadeh

摘要

Sentiment analysis is the process of determining the expressive direction of the user reviews. Recently, sentiment analysis gets more attention. However, low data sentiment analysis receives less attention. The existing works try to augment the samples to consider this issue. In this study, we have utilized a semi-supervised approach to propose a new approach for low-data sentiment analysis. To do so, we have utilized pre-trained XLNet as a feature extractor network to initialize the feature vector for each tweet. Next, these initial representations are fed into the embedding update module to map features into the new space by optimizing the contrastive loss. Then, we utilized a semi-supervised boosting method to assign pseudo labels to unlabeled data. The iteration between the semi-supervised module and the embedding update module is done until convergence is happened. During these iterations, the embedding update module propagates the error-correcting signals to a semi-supervised module. To evaluate the proposed approach, we have applied it to the SemEval2017dataset (task 4), Sentiment 140, and IMDB Movie Reviews. We have designed many different experiment settings to validate the proposed approach’s different modules. On SemEval2017dataset (task 4), we have got 75.9% and 77.1% in AvgRec and \({F}_{1}^{PN}\) F 1 PN respectively. Also, when only 10% of the training samples as labeled samples are used, we get the 71.8% and 73.6% in AvgRec and \({F}_{1}^{PN}\) F 1 PN respectively. The results show that our approach significantly improves with respect to the comparable methods. Also, on IMDB Movie Reviews and Sentiment 140, the proposed approach demonstrates improved performance compared to comparable methods.