Emotional Features in Speech: A Deep Scattering-Based Analysis
摘要
Speech Emotion Recognition (SER), which enables systems to identify emotional states from speech, plays a crucial role in advancing Human-Computer Interaction (HCI). This study explores the integration of Wavelet Scattering Transform (WST), a robust feature extraction technique that captures multiscale temporal-spectral information, with five deep learning models: CNN, SqueezeNet, ResNet, AlexNet, and Self-Attention CNN. Using the TESS dataset, we evaluate the effectiveness of WST in enhancing SER performance. Among the models, CNN and Self-Attention CNN achieved superior accuracy in emotion classification, demonstrating the potential of WST-based representations for SER. Our findings highlight the comparative strengths of these approaches and their relevance to emotion recognition tasks.