<p>Human thoughts, feelings, and ideas are expressed through speech. Stuttering, also referred to as stammering, is an impediment to speech that affects millions of people across the globe. Stuttering speech recognition is a good deal of research within the realm of speech signal processing. Classification of eight different types of stuttering including fluent, interjections, broken words, prolongations, word repetitions, phrase–repetitions, part-word repetitions, and sound repetitions is the main motivation in this research. Stuttering speech recognition is becoming easier with the advancement of machine learning and deep learning. The present study focuses on the efficacy of stuttering speech recognition through the application of machine learning &amp; deep learning. Analysis in this study began with traditional machine learning algorithms, such as Decision Tree (DT), Random Forest (RF), K-Nearest Neighbor (KNN), and Support Vector Machine (SVM). It is seen, feature extraction is an essential component in achieving accurate classification for machine learning algorithms. MFCC &amp; MEL spectrogram features are extracted for this experiment. Furthermore, Delta2 MFCC and Delta2 Mel Spectrogram are derived by second-order derivatives on MFCC &amp; Mel Spectrogram features. Therefore, a variety of these features are employed as input to the model. In comparison with all the aforesaid traditional machine learning models, the KNN model results in the highest accuracy of 69.1% based on MFCC-40 input features. To improve accuracy, we move to a deep learning approach. In this experiment, the convolutional neural network (CNN) model produces better outcomes than traditional machine learning. The study presents a comprehensive analysis of the CNN model with MFCC &amp; MEL spectrogram features, as well as its efficiency. It is noted that MFCC-40 input features in CNN model results in the highest accuracy of 89% after 400 epochs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep analysis of MFCC and MEL spectrogram features to recognize and classify stuttered speech

  • Nilanjan Banerjee,
  • Nilambar Sethi,
  • Samarjeet Borah

摘要

Human thoughts, feelings, and ideas are expressed through speech. Stuttering, also referred to as stammering, is an impediment to speech that affects millions of people across the globe. Stuttering speech recognition is a good deal of research within the realm of speech signal processing. Classification of eight different types of stuttering including fluent, interjections, broken words, prolongations, word repetitions, phrase–repetitions, part-word repetitions, and sound repetitions is the main motivation in this research. Stuttering speech recognition is becoming easier with the advancement of machine learning and deep learning. The present study focuses on the efficacy of stuttering speech recognition through the application of machine learning & deep learning. Analysis in this study began with traditional machine learning algorithms, such as Decision Tree (DT), Random Forest (RF), K-Nearest Neighbor (KNN), and Support Vector Machine (SVM). It is seen, feature extraction is an essential component in achieving accurate classification for machine learning algorithms. MFCC & MEL spectrogram features are extracted for this experiment. Furthermore, Delta2 MFCC and Delta2 Mel Spectrogram are derived by second-order derivatives on MFCC & Mel Spectrogram features. Therefore, a variety of these features are employed as input to the model. In comparison with all the aforesaid traditional machine learning models, the KNN model results in the highest accuracy of 69.1% based on MFCC-40 input features. To improve accuracy, we move to a deep learning approach. In this experiment, the convolutional neural network (CNN) model produces better outcomes than traditional machine learning. The study presents a comprehensive analysis of the CNN model with MFCC & MEL spectrogram features, as well as its efficiency. It is noted that MFCC-40 input features in CNN model results in the highest accuracy of 89% after 400 epochs.