Analysis of Speech Recognition and Pronunciation Evaluation System in English Education Based on Deep Learning and Acoustic Modeling
摘要
Computer-assisted language learning technology may provide a solution to this problem, given the progress in both computer science and technology and in the ways of teaching and learning languages. Speech recognition technology, together with other evaluation tools, is the backbone of computer-assisted language learning. Deep learning is an effective technique for learning the fundamental aspects of a dataset from a small sample size since it approximates complicated functions by learning deep nonlinear network topologies. Using Mel-frequency cepstral coefficient characteristics derived from human auditory models and deep belief networks, this study uses deep learning technology to English voice recognition. By comparing the model’s findings to those of enhanced hidden Markov models, BP neural network models, and tree-based approximation models, the Spoken ArabicDigit dataset from the UCI Machine Learning Repository is used to verify the model’s recognition performance. This research seeks to advance previous computer-based English pronunciation quality rating techniques by focusing on the speech of university students. Pitch precision, speech tempo, rhythm, and intonation are only few of the factors used in the evaluation process. Intonation is evaluated on the basis of pitch frequency, pitch accuracy is evaluated using Mel-frequency cepstral coefficient characteristics, speech pace is evaluated using speech duration, rhythm is evaluated using short-time energy and pairwise variation index, and so on. The evaluation procedures for pitch accuracy, speech pace, rhythm, and intonation utilized in this study have been shown to be reliable by experimental data.