AugMixSpeech: A Data Augmentation Method and Consistency Regularization for Mandarin Automatic Speech Recognition
摘要
Automatic speech recognition (ASR) is a crucial technology in the field of artificial intelligence, widely applied in modern society. The deep learning-based ASR method offers a simpler training framework and higher recognition rates compared to the traditional method. However, it requires large amounts of training data to perform well, and insufficient data can lead to model overfitting. To overcome these problems, we propose a novel data augmentation framework called AugMixSpeech, which generates more natural and diverse data by randomly sampling different augmentation techniques and mixing the augmented data. Besides, in order to ensure that the model maintains stable predictions when faced with these data, we introduce a consistency regularization method that includes global consistency and local consistency. The constraints imposed by this method enable the model to better learn the intrinsic features of the data. Extensive experiments on the validation and test of Aishell-1 achieve recognition accuracy of 4.23% and 4.79%, which outperforms existing approaches and demonstrates its effectiveness in automatic speech recognition.