Along with computer vision algorithms for processing photo and video information, as well as natural language techniques for text analysis, working with audio information is also the most demanded procedure for conducting business analytics. The article considers the problem of speech signal recognition using the example of the audio database formed on the basis of Azerbaijani words reproduced by native speakers. In the proposed approach, the sound signal is considered as a one-dimensional representation of sound wave oscillations with a certain sampling frequency. To implement the task, the amplitude recognition method, as well as the Dynamic Time Warping and Derivative Dynamic Time Warping methods, which have proven themselves well in the subject area of the study, are used. Empirical analysis of the results of recognizing voice reproductions of words in the Azerbaijani language revealed the qualitative advantages of the DTW method over the others. In the context of the study, a mechanism for forming an audio database is proposed based on the iterative method for clarifying the reliability of a voice signal reflecting a homograph word voiced by different native speakers in different pronunciations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Formation of a Language Model for Recognizing Spoken Words from the Azerbaijani Lexicon

  • Ramin Rzayev,
  • Elchin Ismayilov,
  • Azer Kerimov

摘要

Along with computer vision algorithms for processing photo and video information, as well as natural language techniques for text analysis, working with audio information is also the most demanded procedure for conducting business analytics. The article considers the problem of speech signal recognition using the example of the audio database formed on the basis of Azerbaijani words reproduced by native speakers. In the proposed approach, the sound signal is considered as a one-dimensional representation of sound wave oscillations with a certain sampling frequency. To implement the task, the amplitude recognition method, as well as the Dynamic Time Warping and Derivative Dynamic Time Warping methods, which have proven themselves well in the subject area of the study, are used. Empirical analysis of the results of recognizing voice reproductions of words in the Azerbaijani language revealed the qualitative advantages of the DTW method over the others. In the context of the study, a mechanism for forming an audio database is proposed based on the iterative method for clarifying the reliability of a voice signal reflecting a homograph word voiced by different native speakers in different pronunciations.