Formation of a Language Model for Recognizing Spoken Words from the Azerbaijani Lexicon
摘要
Along with computer vision algorithms for processing photo and video information, as well as natural language techniques for text analysis, working with audio information is also the most demanded procedure for conducting business analytics. The article considers the problem of speech signal recognition using the example of the audio database formed on the basis of Azerbaijani words reproduced by native speakers. In the proposed approach, the sound signal is considered as a one-dimensional representation of sound wave oscillations with a certain sampling frequency. To implement the task, the amplitude recognition method, as well as the Dynamic Time Warping and Derivative Dynamic Time Warping methods, which have proven themselves well in the subject area of the study, are used. Empirical analysis of the results of recognizing voice reproductions of words in the Azerbaijani language revealed the qualitative advantages of the DTW method over the others. In the context of the study, a mechanism for forming an audio database is proposed based on the iterative method for clarifying the reliability of a voice signal reflecting a homograph word voiced by different native speakers in different pronunciations.