Designing expressive speech synthesis for child voice remains an unresolved problem. One of the major dilemmas faced by child TTS systems and child speech synthesis is the scarcity of datasets to train opaque data-hungry DNN-based models. Only a few datasets were proposed for the purpose of building child conversational AI agents, and many of them come with challenges such as noisy data and indiscernible speech. With this in mind, we introduce the ChildTinyTalks (CTT) dataset, comprising 2 h of speech collected from 25 kids in grades ranging from third to fourth grade, who are telling stories and sharing their experiences. The new dataset containing 1200 audio samples has been transcribed at the word level, comprising 4 classes of voice expressions. To verify the effectiveness of CTT in real-world situations, AutoVocoder models were trained and synthesized samples were generated. The models were trained on both the LJSpeech large scale dataset and our CTT dataset. Initial experimental results indicate that the CTT dataset can steadily give comparable results with acoustic model trained on a large-scale dataset with a size of less than 10% of the large dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ChildTinyTalks (CTT): A Benchmark Dataset and Baseline for Expressive Child Speech Synthesis

  • Shaimaa Alwaisi,
  • Mohammed Salah Al-Radhi,
  • Géza Németh

摘要

Designing expressive speech synthesis for child voice remains an unresolved problem. One of the major dilemmas faced by child TTS systems and child speech synthesis is the scarcity of datasets to train opaque data-hungry DNN-based models. Only a few datasets were proposed for the purpose of building child conversational AI agents, and many of them come with challenges such as noisy data and indiscernible speech. With this in mind, we introduce the ChildTinyTalks (CTT) dataset, comprising 2 h of speech collected from 25 kids in grades ranging from third to fourth grade, who are telling stories and sharing their experiences. The new dataset containing 1200 audio samples has been transcribed at the word level, comprising 4 classes of voice expressions. To verify the effectiveness of CTT in real-world situations, AutoVocoder models were trained and synthesized samples were generated. The models were trained on both the LJSpeech large scale dataset and our CTT dataset. Initial experimental results indicate that the CTT dataset can steadily give comparable results with acoustic model trained on a large-scale dataset with a size of less than 10% of the large dataset.