<p>In the contemporary realm of speech emotion recognition (SER), the synergy of human speech attributes and advanced deep learning (DL) techniques has significantly improved the emotion recognition performance. This study presents a pioneering contribution to SER by employing speech representations extracted from the self-supervised learning framework WavLM. These representations are coupled with a specialized statistical feature process designed to reduce the high-dimensional feature space. Furthermore, a DL architecture named Minimal-auto-encoder (MiAE) is developed to compress the embeddings while preserving the most relevant and discriminative representations. These compressed features are subsequently fed into the classification layer to ensure accurate emotion recognition. The proposed SER methodology has undergone rigorous evaluation through a series of experiments conducted in both speaker dependent (SD) and speaker independent (SI) scenarios. The comparative analysis involves existing SER techniques, underscoring the superiority of the proposed approach in terms of established performance metrics. Notably, the experimental outcomes, derived from three distinct datasets namely IEMOCAP, RAVDESS, and SAVEE reveal remarkable accuracies of <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(76.49\%\)</EquationSource> </InlineEquation>, <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(84.72\%\)</EquationSource> </InlineEquation>, and <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(87.50\%\)</EquationSource> </InlineEquation> respectively in the SD scenario. In the SI scenario, corresponding results stand at <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(69.03\%\)</EquationSource> </InlineEquation>, <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(72.29\%\)</EquationSource> </InlineEquation>, and <InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(62.90\%\)</EquationSource> </InlineEquation> respectively. These findings substantiate the efficacy and potential of the proposed MiAE-based SER method in advancing the state-of-the-art.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing speech emotion recognition using statistical self-supervised embeddings

  • Adil Chakhtouna,
  • Sara Sekkate,
  • Abdellah Adib

摘要

In the contemporary realm of speech emotion recognition (SER), the synergy of human speech attributes and advanced deep learning (DL) techniques has significantly improved the emotion recognition performance. This study presents a pioneering contribution to SER by employing speech representations extracted from the self-supervised learning framework WavLM. These representations are coupled with a specialized statistical feature process designed to reduce the high-dimensional feature space. Furthermore, a DL architecture named Minimal-auto-encoder (MiAE) is developed to compress the embeddings while preserving the most relevant and discriminative representations. These compressed features are subsequently fed into the classification layer to ensure accurate emotion recognition. The proposed SER methodology has undergone rigorous evaluation through a series of experiments conducted in both speaker dependent (SD) and speaker independent (SI) scenarios. The comparative analysis involves existing SER techniques, underscoring the superiority of the proposed approach in terms of established performance metrics. Notably, the experimental outcomes, derived from three distinct datasets namely IEMOCAP, RAVDESS, and SAVEE reveal remarkable accuracies of \(76.49\%\) , \(84.72\%\) , and \(87.50\%\) respectively in the SD scenario. In the SI scenario, corresponding results stand at \(69.03\%\) , \(72.29\%\) , and \(62.90\%\) respectively. These findings substantiate the efficacy and potential of the proposed MiAE-based SER method in advancing the state-of-the-art.