Abstract <p>A study of the influence of differences in languages and training data on the quality of cross-lingual transfer of a trained speech model to Russian in the task of automatic recognition of emotions in speech is described. At the training stage, English, Polish, Chinese, and Japanese served as source languages, for which the IEMOCAP, nEMO, ESD, and JVNV emotional speech datasets were used, respectively, and the model itself was the HuBERT speech model on the transformer architecture. All models trained on the corresponding dataset were tested on a shortened sample from the Dusha Russian emotional speech dataset. Based on the data obtained, the main trends in choosing different languages for training the speech model and its subsequent transfer to Russian are considered, and differences in datasets are analyzed, which indicate the need for further work on collecting and labeling quality emotional speech data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Lingual Transfer for Russian Speech Emotion Automatic Recognition: Data and Trends

  • V. I. Lemaev,
  • N. V. Lukashevich

摘要

Abstract

A study of the influence of differences in languages and training data on the quality of cross-lingual transfer of a trained speech model to Russian in the task of automatic recognition of emotions in speech is described. At the training stage, English, Polish, Chinese, and Japanese served as source languages, for which the IEMOCAP, nEMO, ESD, and JVNV emotional speech datasets were used, respectively, and the model itself was the HuBERT speech model on the transformer architecture. All models trained on the corresponding dataset were tested on a shortened sample from the Dusha Russian emotional speech dataset. Based on the data obtained, the main trends in choosing different languages for training the speech model and its subsequent transfer to Russian are considered, and differences in datasets are analyzed, which indicate the need for further work on collecting and labeling quality emotional speech data.