<p>Speech-based sentiment recognition (SSR) is imperative in many human–machine interaction systems. Distinct deep learning (DL)-based techniques have been investigated for SSR in the last decade, which have shown meaningful enhancement over machine learning (ML)-based SSR schemes. However, the outcomes of these techniques are limited because of the poor feature representation, lower long-term dependency, redundancy in features, and poor intra-class and interclass disparity. This article presents the SSR based on a parallel deep convolutional neural network and long short term memory (PDCNN-LSTM). The PDCNN-LSTM considers multiple spectral-domain voice features (SDVF), time-domain voice features (TDVF), and voice-quality features (VQF). Furthermore, it uses Archimedes' optimization algorithm (AoA) for feature selection. The multi-attribute utility theory (MAUT) technique automatically decides the weights of the AoA Objective function. The PDCNN-LSTM provides improved accuracy of 99.22%, precision of 0.99, recall of 0.99, and F1-score of 0.99 compared to existing techniques.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speech sentiment recognition using lightweight parallel deep convolution neural network and long short-term memory

  • Sushadevi Shamrao Adagale,
  • Praveen Gupta,
  • Shubhangi Adagale-Vairagar

摘要

Speech-based sentiment recognition (SSR) is imperative in many human–machine interaction systems. Distinct deep learning (DL)-based techniques have been investigated for SSR in the last decade, which have shown meaningful enhancement over machine learning (ML)-based SSR schemes. However, the outcomes of these techniques are limited because of the poor feature representation, lower long-term dependency, redundancy in features, and poor intra-class and interclass disparity. This article presents the SSR based on a parallel deep convolutional neural network and long short term memory (PDCNN-LSTM). The PDCNN-LSTM considers multiple spectral-domain voice features (SDVF), time-domain voice features (TDVF), and voice-quality features (VQF). Furthermore, it uses Archimedes' optimization algorithm (AoA) for feature selection. The multi-attribute utility theory (MAUT) technique automatically decides the weights of the AoA Objective function. The PDCNN-LSTM provides improved accuracy of 99.22%, precision of 0.99, recall of 0.99, and F1-score of 0.99 compared to existing techniques.