The landscape of speech synthesis technology, particularly neural Text-to-Speech (TTS), has seen rapid advancements in recent years. This review examines the current state of neural TTS systems and their availability for edge on-device deployment. Traditionally, neural TTS models have required substantial computational resources, limiting their application to server-based cloud implementations. However, recent innovations in model architecture and synthesis techniques are making it possible to deploy these systems on edge devices with limited processing power. These developments are crucial for applications requiring low latency, in which it is necessary that data is processed locally without reliance on cloud services. One important use-case is in assistive technology such as Augmented and Alternative Communication (AAC) for users with speech impairments and screen readers for the visually impaired.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Review of the State of the Speech Synthesis Technology Landscape – Neural TTS on the Edge

  • Branislav Gerazov,
  • Vanesa Lazareva,
  • Marija Markovska Dimitrovska,
  • Dimitar Taskovski,
  • Katerina Mavrou,
  • Eleni Theodorou,
  • Francesco Zanfardino,
  • Antonio Spera,
  • Antonello Mura,
  • Anna Rybińska,
  • May Agius,
  • Danche Todorovska,
  • Nefi Charalambous-Darden,
  • Katarzyna Łuszczak,
  • Silvio Pagliara

摘要

The landscape of speech synthesis technology, particularly neural Text-to-Speech (TTS), has seen rapid advancements in recent years. This review examines the current state of neural TTS systems and their availability for edge on-device deployment. Traditionally, neural TTS models have required substantial computational resources, limiting their application to server-based cloud implementations. However, recent innovations in model architecture and synthesis techniques are making it possible to deploy these systems on edge devices with limited processing power. These developments are crucial for applications requiring low latency, in which it is necessary that data is processed locally without reliance on cloud services. One important use-case is in assistive technology such as Augmented and Alternative Communication (AAC) for users with speech impairments and screen readers for the visually impaired.