Enhancing speech emotion recognition using statistical self-supervised embeddings
摘要
In the contemporary realm of speech emotion recognition (SER), the synergy of human speech attributes and advanced deep learning (DL) techniques has significantly improved the emotion recognition performance. This study presents a pioneering contribution to SER by employing speech representations extracted from the self-supervised learning framework WavLM. These representations are coupled with a specialized statistical feature process designed to reduce the high-dimensional feature space. Furthermore, a DL architecture named Minimal-auto-encoder (MiAE) is developed to compress the embeddings while preserving the most relevant and discriminative representations. These compressed features are subsequently fed into the classification layer to ensure accurate emotion recognition. The proposed SER methodology has undergone rigorous evaluation through a series of experiments conducted in both speaker dependent (SD) and speaker independent (SI) scenarios. The comparative analysis involves existing SER techniques, underscoring the superiority of the proposed approach in terms of established performance metrics. Notably, the experimental outcomes, derived from three distinct datasets namely IEMOCAP, RAVDESS, and SAVEE reveal remarkable accuracies of