Automatic speech recognition (ASR) is a crucial technology in the field of artificial intelligence, widely applied in modern society. The deep learning-based ASR method offers a simpler training framework and higher recognition rates compared to the traditional method. However, it requires large amounts of training data to perform well, and insufficient data can lead to model overfitting. To overcome these problems, we propose a novel data augmentation framework called AugMixSpeech, which generates more natural and diverse data by randomly sampling different augmentation techniques and mixing the augmented data. Besides, in order to ensure that the model maintains stable predictions when faced with these data, we introduce a consistency regularization method that includes global consistency and local consistency. The constraints imposed by this method enable the model to better learn the intrinsic features of the data. Extensive experiments on the validation and test of Aishell-1 achieve recognition accuracy of 4.23% and 4.79%, which outperforms existing approaches and demonstrates its effectiveness in automatic speech recognition.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AugMixSpeech: A Data Augmentation Method and Consistency Regularization for Mandarin Automatic Speech Recognition

  • Yang Jiang,
  • Jun Chen,
  • Kai Han,
  • Yi Liu,
  • Siqi Ma,
  • Yuqing Song,
  • Zhe Liu

摘要

Automatic speech recognition (ASR) is a crucial technology in the field of artificial intelligence, widely applied in modern society. The deep learning-based ASR method offers a simpler training framework and higher recognition rates compared to the traditional method. However, it requires large amounts of training data to perform well, and insufficient data can lead to model overfitting. To overcome these problems, we propose a novel data augmentation framework called AugMixSpeech, which generates more natural and diverse data by randomly sampling different augmentation techniques and mixing the augmented data. Besides, in order to ensure that the model maintains stable predictions when faced with these data, we introduce a consistency regularization method that includes global consistency and local consistency. The constraints imposed by this method enable the model to better learn the intrinsic features of the data. Extensive experiments on the validation and test of Aishell-1 achieve recognition accuracy of 4.23% and 4.79%, which outperforms existing approaches and demonstrates its effectiveness in automatic speech recognition.