Classification of Speech Signal Using CNN-LSTM
摘要
Speech categorisation is a broad study that has attracted a lot of interest recently. Understanding other people's emotions and responding appropriately is the largest distinction between robots and people. Speech signals are categorised in this article using both conventional machine learning methods and deep learning algorithms. Speech emotion is categorised using the interactive emotional binary motion capture (IEMOCAP) and Ryerson affective voice and audio-visual database of songs (RAVDESS) databases. The expanded dataset is mixed with the voice signals once they have been transformed into spectrograms. This model incorporates CNN + LSTM, attention, and a novel type of spectrogram frequency distribution based on the spectrogram’s properties. Mel-frequency cepstral coefficients guarantee improved speech feature performance in tasks requiring emotion categorisation.