Child voices in Text-To-Speech (TTS) are essential for enabling children with speech or communication difficulties to express themselves authentically, supporting their social inclusion and sense of identity. However, the development and availability of high-quality child voices remain limited compared to adult voices, mostly because of the lack of child data and ethical reasons considering that the voice is owned by a minor. In this paper we review the current state of the TTS synthesis technological landscape with special focus on systems that are offering child voices, as well as technology that is targeting offline use on low resource mobile devices, such as smartphones and tablets. We also explore research efforts in creating child TTS by reviewing papers on neural child TTS models in terms of the technologies that are used as well as the quality of voices. Additionally, we make a short summary of the available TTS engines and their voices to check the availability of child voices. Finally, we examine the use of child voices in available tools and applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Review of the State of the Child Speech Synthesis Landscape

  • Vanesa Lazareva,
  • Marija Markovska Dimitrovska,
  • Silvio Pagliara,
  • Katerina Mavrou,
  • Eleni Theodorou,
  • Dimitar Taskovski,
  • Danche Todorovska,
  • Francesco Zanfardino,
  • Antonio Spera,
  • Antonello Mura,
  • Anna Rybińska,
  • May Agius,
  • Nefi Charalambous-Darden,
  • Katarzyna Łuszczak,
  • Branislav Gerazov

摘要

Child voices in Text-To-Speech (TTS) are essential for enabling children with speech or communication difficulties to express themselves authentically, supporting their social inclusion and sense of identity. However, the development and availability of high-quality child voices remain limited compared to adult voices, mostly because of the lack of child data and ethical reasons considering that the voice is owned by a minor. In this paper we review the current state of the TTS synthesis technological landscape with special focus on systems that are offering child voices, as well as technology that is targeting offline use on low resource mobile devices, such as smartphones and tablets. We also explore research efforts in creating child TTS by reviewing papers on neural child TTS models in terms of the technologies that are used as well as the quality of voices. Additionally, we make a short summary of the available TTS engines and their voices to check the availability of child voices. Finally, we examine the use of child voices in available tools and applications.