<p>Verbal communication remains the most widely used form of interaction, and the development of speech synthesis that accurately conveys emotion is an increasingly important area of research in speech processing. This is especially relevant for applications in voice assistants, robotics, e-learning, and assistive technologies for individuals with disabilities. Speech emotion synthesis involves generating synthetic speech that reflects various emotional states such as happiness, sadness, anger, and excitement. This process typically combines techniques from speech synthesis (e.g., text-to-speech or TTS) with advanced speech information processing. In this work, we introduce a novel approach that integrates a speech synthesis model with a speech context generation module powered by a Large Language Model (LLM) to predict and embed underlying emotional cues. We further enhance the TacotronDDC model by replacing the traditional vocoder with a context-ingesting module that incorporates emotion-related metadata derived from the LLM attaining an f1 score of 81% in emotion prediction as well as MOS and PESQ score of 3.92 and 1.95 in synthesis respectively. Several challenges remaining in this domain, such as addressing the context-dependent nature of emotions, integrating multimodal emotion recognition, and managing the limitations of imbalanced and small datasets. To evaluate our approach, an open-loop prediction tests using the RAVDESS and IEMOCAP datasets has been conducted. The results show that our framework effectively leverages the complementary knowledge from different modules, outperforming existing baseline synthesis methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EmoSRE: Emotion prediction based speech synthesis and refined speech recognition using large language model and prosody encoding

  • Shivam Akhouri,
  • Ananthakrishnan Balasundaram

摘要

Verbal communication remains the most widely used form of interaction, and the development of speech synthesis that accurately conveys emotion is an increasingly important area of research in speech processing. This is especially relevant for applications in voice assistants, robotics, e-learning, and assistive technologies for individuals with disabilities. Speech emotion synthesis involves generating synthetic speech that reflects various emotional states such as happiness, sadness, anger, and excitement. This process typically combines techniques from speech synthesis (e.g., text-to-speech or TTS) with advanced speech information processing. In this work, we introduce a novel approach that integrates a speech synthesis model with a speech context generation module powered by a Large Language Model (LLM) to predict and embed underlying emotional cues. We further enhance the TacotronDDC model by replacing the traditional vocoder with a context-ingesting module that incorporates emotion-related metadata derived from the LLM attaining an f1 score of 81% in emotion prediction as well as MOS and PESQ score of 3.92 and 1.95 in synthesis respectively. Several challenges remaining in this domain, such as addressing the context-dependent nature of emotions, integrating multimodal emotion recognition, and managing the limitations of imbalanced and small datasets. To evaluate our approach, an open-loop prediction tests using the RAVDESS and IEMOCAP datasets has been conducted. The results show that our framework effectively leverages the complementary knowledge from different modules, outperforming existing baseline synthesis methods.