With the advancement of information technology, Computer-Assisted Pronunciation Training (CAPT) has become an effective method for non-native(L2) speakers to learn foreign language pronunciation. However, existing automatic pronunciation quality assessment methods have not fully leveraged the inter-granularity relationships and lack further extraction of contextual features at each granularity. To address these issues, this paper proposes Bfhaformer. Bfhaformer employs an LSTM-augmented BranchFormer encoder for encoding GOP features and reference phoneme features. Compared to Transformer encoders, the BranchFormer encoder introduces parallel branch structures, which enhances the capture of local features while retaining global feature information. Additionally, this paper aggregates features across different granularities within a hierarchical model structure. By aggregating and suprasegmental feature fusion of the encoded features at pronunciation granularity such as word level and utterance level, better attention is paid to local information at the current granularity and contextual hierarchical relationships. Experiments on the publicly available Speechocean762 dataset demonstrate that our proposed method significantly improves all metrics at all granularities compared to the baseline models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multi-aspect Multi-granularity Pronunciation Assessment Method Based on Branchformer Encoder and Hierarchical Aggregation

  • Wenxu Du,
  • Aishan Wumaier,
  • Yahui Shi,
  • Nian Yi,
  • Dehua Liu

摘要

With the advancement of information technology, Computer-Assisted Pronunciation Training (CAPT) has become an effective method for non-native(L2) speakers to learn foreign language pronunciation. However, existing automatic pronunciation quality assessment methods have not fully leveraged the inter-granularity relationships and lack further extraction of contextual features at each granularity. To address these issues, this paper proposes Bfhaformer. Bfhaformer employs an LSTM-augmented BranchFormer encoder for encoding GOP features and reference phoneme features. Compared to Transformer encoders, the BranchFormer encoder introduces parallel branch structures, which enhances the capture of local features while retaining global feature information. Additionally, this paper aggregates features across different granularities within a hierarchical model structure. By aggregating and suprasegmental feature fusion of the encoded features at pronunciation granularity such as word level and utterance level, better attention is paid to local information at the current granularity and contextual hierarchical relationships. Experiments on the publicly available Speechocean762 dataset demonstrate that our proposed method significantly improves all metrics at all granularities compared to the baseline models.