Empirical Analysis of Conformer Impact on the CTC-CRF Model in Kazakh Speech Recognition
摘要
The enhancement of quality and efficiency in contemporary speech technology systems is increasingly reliant on the application of machine learning techniques. Recent advancements have highlighted the superior performance of end-to-end speech recognition methods. Integrating multiple end-to-end model architectures within a single network has yielded competitive outcomes. This study focuses on the development and evaluation of an integrated Connectionist Temporal Classification-Conditional Random Field system, augmented with the Conformer, to enhance the recognition accuracy of a low-resource language such as Kazakh. Due to the limited electronic and digital resources in the Kazakh language, extensive efforts are required for the collection and compilation of a comprehensive speech and text corpus. Additionally, this necessitates the development of novel mathematical models and algorithms to address the challenges in automatic speech recognition for agglutinative (Turkic) languages. The research findings demonstrate that the incorporation of the Conformer model alongside CTC-CRF frameworks significantly improves speech recognition outcomes. The system, utilizing the ResNet network and RNN-based language models, further enriched with the Conformer, attained optimal performance, enhancing the recognition accuracy by 1.6%.